Homa Claims 13x Better Tail Latency Than TCP for AI Clusters. Read the Conditions First
Hardware / explainer
Homa Claims 13x Better Tail Latency Than TCP for AI Clusters. Read the Conditions First
Stanford's John Ousterhout told the AI Engineer World's Fair that TCP and RoCE suit bulk transfers, not millisecond-scale AI traffic. His own caveats are the useful part.

Homa, a message-based transport protocol from Stanford's John Ousterhout, cut 99th-percentile latency for short messages by roughly 13 times against TCP in the results he presented: under 100 microseconds versus over 1 millisecond. The same results show only about a 2x gain for the longest messages, and by the talk's own summary they do not yet show an end-to-end application speedup or a comparison with RoCE.
Ousterhout presented the work at the AI Engineer World's Fair 2026 in a talk titled "Homa: The End of TCP for AI Clusters". The video drew 44 points on Hacker News. The headline number is a tail-latency figure, so the rest of this piece is about what it measures.
Why short messages decide the argument
The talk's thesis is that AI traffic is shifting from large bulk transfers to a mix that includes many small messages: tool calls, key-value lookups, synchronisation. Ousterhout's line: "Short-message tail latency matters most when compute phases themselves approach the millisecond scale," because GPUs sit idle until every required exchange completes.
If a step takes a millisecond of compute and the slowest short message in the step takes another millisecond to arrive, utilisation halves for that step. That is the denominator behind the 13x figure, and it is why P99 rather than the mean is the quantity that matters. A cluster waits for its slowest message, not its average one.
What TCP and RoCE get wrong, per the talk
The talk names two faults in TCP and RoCE:
- Sender-driven congestion control: senders react only after queues build at a distant receiver, and "it typically takes several round trips for the sender to gradually adjust its rate."
- A byte-stream model: without visible message boundaries, short messages get trapped behind long ones, the head-of-line blocking problem.
Homa is built on three mechanisms: messages as the unit, so "as soon as a receiver gets the first packet of a message, it knows exactly how much more data the sender wants to send"; receiver-issued grants that pace the scheduled portion; and switch priority queues so short messages bypass queued long ones.
The numbers, with their limits
| Measure | Result in the talk | Caveat |
|---|---|---|
| P99, short messages | under 100 microseconds vs over 1 ms on TCP | no RoCE comparison |
| Longest messages | nearly 2x better | not an application speedup |
| homa_qdisc beside TCP | 4x lower P99 for short messages (README, January 2026) | different test from the 13x |
The last row comes from the project's GitHub README, which reports that the homa_qdisc queuing discipline reduced short-message P99 by 4x when Homa ran alongside TCP. It is a separate benchmark, and the two figures should not be stacked.
The README also makes a stronger claim, that tail latency beats TCP by 10x across measured workloads and that Homa's 99th-percentile latency "is usually better than TCP's mean latency." Those are the project's own measurements. The README's list of supported hardware is narrow: Mellanox ConnectX-4, 5 and 6, and Intel E810 NICs.
Where it stands as software
Homa ships as a Linux kernel module. The README lists support through kernel 6.17.8 as of November 2025, backports for Red Hat Enterprise Linux 8 and 9.5, IPv6, network namespaces and gRPC integration. The repository showed 510 stars, 74 forks and protocol changes as recent as September 2026. Upstreaming into the Linux kernel has been under way since October 2024, so it is not in a stock kernel.
Two gaps matter for an operator. The incast optimisation from the project's SIGCOMM paper remains unimplemented, per the README, and incast, many senders hitting one receiver, is a pattern a fan-out of requests to one node would plausibly create (our inference, not the README's).

The adapter above is an InfiniBand card, a sibling of the RDMA hardware the talk argues against, not a Homa-specific device.
What would change this read
KRIT HUB's write-up of the talk offers the sober takeaway: engineers need not replace TCP now, but should profile inference clusters for short-message dominance and measure head-of-line blocking under mixed workloads. The variable that would falsify the case for Homa is simple. If a real serving cluster's step time is dominated by compute or by large collective transfers, short-message tail latency is not the bottleneck and a 13x gain on it buys little.
What would strengthen it: an end-to-end tokens-per-second result on a named model, a RoCE comparison, and a Homa-capable NIC list wider than two vendors. Related reading on this site: Cloudflare's Clef decision models and Intel's tokens-per-watt GPU.
Sources
More in Hardware
- 01YMX Signs Five-Year Deal to Run Outrider's Self-Driving Yard Trucks, With No Fleet Size GivenThe logistics operator will put Outrider's autonomy on its Orange EV electric trucks in customer yards from 2026. The announcement gives no truck count, price or site.
- 02Valve's Timur Kristóf Got a 14-Year-Old Radeon HD 7870 XT Working on LinuxA Valve driver engineer's patches moved Radeon HD 7000 and R9 200 cards to amdgpu by default in Linux 6.19, with a reported 30% uplift and a fix for a card that never worked.
- 03Google's Lincoln Data Center Peaks at 52.65 MW, a Redaction Error ShowsGoogle called the figures trade secrets. Copying text from under the black boxes in a Nebraska state filing produced them, along with $117.6 million in expected tax refunds.
- 04Tesla's Austin Robotaxi Now Closes at 11 PM, an Hour Earlier Than at Launch, as Rivian Prices Lidar at a Few Hundred DollarsElon Musk blamed grey kittens on grey tarmac for the limited night hours. Rivian's lidar cost estimate is the number that makes that explanation awkward for a camera-only fleet.