LiteLLM's Rust Gateway Claims 15x Throughput, but Its Docs List Four Routes and No Benchmarks
Software / analysis
LiteLLM's Rust Gateway Claims 15x Throughput, but Its Docs List Four Routes and No Benchmarks
A June 22 blog post promised OCR, chat and a router by mid-September. The beta docs page read on October 10 shows a narrower path and a Docker image that is not published.
LiteLLM's Rust gateway benchmarks at 6,782 requests a second against 453 for the Python proxy, according to the project's own June 22 blog post. The docs page for the same Rust path, read on October 10, lists four supported routes, says streaming chat requests fall back to Python, and carries no benchmark numbers at all.
The gap between the two pages is the story. The launch post by Ishaan Jaffer, who is listed as LiteLLM's CTO, promised dates. The Rust gateway docs describe what shipped, labelled beta and off by default.
The June 22 numbers, and what they measured
The post reports three results from a benchmark the LiteLLM team ran. The hardware is not specified.
- Rust gateway6782 requests per second
- LiteLLM Python453 requests per second
Source: LiteLLM blog post 'Migrating LiteLLM to Rust', June 22, 2026; vendor-run, hardware not stated
Per-request overhead at 10 concurrent clients was about 0.05 ms for Rust against about 7.5 ms for Python. Peak memory under load was about 31.7 MB against about 358.9 MB. The post rounds that to 15 times the throughput on 11 times less memory.
The method matters more than the ratios. The post describes a mock upstream, a "thin axum forwarding gateway" on the Rust side, and the Python path called through litellm.acompletion over uvicorn. It says the benchmark covers the forwarding path only, "not a full production workload." A forwarding stub with no auth, routing or spend tracking is not the product users run, so the 15x is a ceiling on what removing Python from the hot path could yield.
The roadmap against the docs
The post gave four target dates. The docs, read on October 10, support some of it and are silent on the rest.
| Post target | Milestone | What the docs page shows |
|---|---|---|
| Aug 15, 2026 | OCR, Mistral first | No OCR row in the Mode 1 route table |
| Sep 1, 2026 | /messages, then /chat/completions | /messages for anthropic and azure_ai from v1.94.0; /chat/completions for anthropic and bedrock "ships in an upcoming release" |
| Sep 15, 2026 | Router, fallbacks, retries | Not mentioned on the page |
| Dec 1, 2026 | Full server, pure Rust | Standalone axum binary has "fewer routes", no published Docker image |
The post called its own plan "not a v2 and not a rewrite" and said each route moves "only after it passes our full parity and end-to-end test suite." Slipping a date to protect parity is consistent with that. It is still a slip the docs do not mention.
What the Rust path accepts today
The docs describe two modes. Mode 1 keeps the Python server for auth, routing and callbacks and sends only provider translation and the network call to Rust, switched on per model with rust: true. A response served by Rust carries an x-litellm-rust: true header, which is the only way to confirm which path handled a request.
The supported list is short: /chat/completions for anthropic and bedrock (Converse), Anthropic /v1/messages for anthropic and azure_ai, audio transcription on bedrock, and Responses API WebSockets on openai.
The /chat/completions conditions are narrow. The Rust path takes non-streaming text conversations only, with just max_tokens, temperature, top_p and stop set, plus top_k on anthropic. Streaming, tool calls, images, response_format, extended thinking, prompt caching and n above 1 all go to Python silently. A developer who set rust: true and sent tool calls would see nothing change except the missing header.
One failure rule is worth knowing. Path selection happens before the provider is called, and a failure after that point "comes back as an error rather than being retried on Python," because a retry would bill twice.
Why Python 3.15 users should read the packaging line
The Rust core ships inside the Python wheel, so there is no separate install. That also ties it to the wheel's constraints. The PyPI metadata for litellm 1.104.2 declares requires_python as <3.15,>=3.10, which excludes the Python 3.15.0 interpreter that python.org released on October 9.
The repository itself is busy. The GitHub releases feed shows v1.106.0-dev.3 published October 9, with notes such as "refactor(rust): rename litellm-core to litellm-inference" and "feat(rust): add standalone typed LLM wire contracts." The renames suggest the crate layout is still moving. The project had 60.7k stars and 1.8k open issues when read, and the package declares an MIT licence expression.
What this means for a gateway you already run
Anyone running the Python proxy loses nothing by waiting, because the Python path stays the default. Anyone who wants the speedup has a narrow, checkable test: enable rust: true on one Anthropic model, send plain non-streaming text, and look for the header. Measuring against your own traffic is the only way to learn whether 7.5 ms of overhead was ever your bottleneck, since the model call itself takes far longer.
The comparison to other rewrites is instructive. This site has reported on GitHub's Rust rewrite of Copilot's runtime and on how long that took against Bun. LiteLLM is attempting a staged version of the same move, and its December 1 full-server date is the next checkpoint against which the September misses can be judged.
Sources
More in Software
- 01Python 3.15.0 Ships October 9 After Lazy-Import Bugs Forced an Unplanned Third CandidateThe release team had called rc2 final on September 1. Late blockers in PEP 810 added an rc3 on October 2 and pushed the stable build to October 9.
- 02Diagram Design Passes 47,900 Stars, and Its README Tells Git Users to Pick MermaidThe agent skill draws 44 diagram types as a single HTML file. Its own README admits layouts vary between runs, and five export bugs were filed on October 6.
- 03Anthropic's Knowledge-Work Plugins Have 28,300 Stars, a Datadog URL Due to Break and a Contribution Guide the CI OverridesThe 11-plugin repo is Apache-2.0 and trending. Its README asks for pull requests, a workflow closes them, and 76 issues sit open.
- 04Unison Cloud Goes MIT, but Its Local-Run Guide Still Pulls Nimbus From a Private RepoPaul Chiusano's Oct. 8 post open-sources the worker node, API server and web UI, while the repo's own instructions show a credentials step the post never mentions.