antirez's ds4 Reports 39.4 Tokens a Second on a 128GB Mac, and Its README Calls It Beta
Software / news
antirez's ds4 Reports 39.4 Tokens a Second on a 128GB Mac, and Its README Calls It Beta
The MIT-licensed C engine from Redis's creator reached 22.9k GitHub stars; its README says AI coding agents wrote much of it.

Salvatore Sanfilippo, the creator of Redis who publishes as antirez, has built an MIT-licensed C inference engine for large open-weight models called DwarfStar 4, or ds4. Its site reports 39.4 tokens a second of generation on a 128GB M5 Max, and the repository had 22.9k stars on October 3.
The project's site and GitHub README describe a purpose-built engine rather than a generic runner. It targets DeepSeek V4 Flash, V4.1 Flash and V4 Pro, plus GLM 5.x and Qwen3.8 Flash Next, on Metal, CUDA and ROCm.

What the speed numbers say
The site lists two machines, both measured at a 2K context. An M5 Max with 128GB reaches 790.2 tokens a second on prefill and 39.4 on generation. An NVIDIA DGX Spark with 128GB reaches 825.8 on prefill and 18.1 on generation.
- Apple M5 Max, 128GB39.4 tokens/s
- NVIDIA DGX Spark, 128GB18.1 tokens/s
Source: dwarfstar.sh, accessed 2026-10-03; project-supplied
These are the author's own figures. The page does not say which model or quantization produced them, so the two rows cannot be read as a ranking of the chips.
For scale, an unrelated arXiv paper by Siming Huang reports about 74.5 tokens a second of decode for DeepSeek-V4-Flash on AMD Instinct MI250 GPUs. It does not mention ds4, and the hardware class is different, so it shows only that other routes to the same model exist.
How a 2-bit model fits
The site's technical claim is "asymmetric 2-bit quantization" applied to the routed experts while the paths the model depends on most keep more precision. The README says these models "tolerate aggressive routed-expert quantization", and it borrows the IQ2_XXS, Q2_K and Q4_K formats from the llama.cpp GGUF ecosystem.
Memory is the constraint. The README names 96GB as the floor for Metal, with smaller machines relying on SSD streaming. DeepSeek V4.1 Flash at Q4 needs a 512GB Mac Studio or streaming, while the Qwen3.8 Flash Q2 build lists 41.73 GiB of main weights and starts at 64GB with an 8K context.
The KV cache can stream to SSD and survives restarts through prompt hashing. Three programs ship: a CLI, ds4-server, which speaks the OpenAI and Anthropic APIs, and a native agent, ds4-agent.
What the README admits
Two sentences carry most of the caveats. The first reads: "The software is currently very fast changing. Consider it beta quality." The second says a large QA run precedes each release, "however instabilities and regressions are definitely possible". Sessions containing images cannot yet be saved.
The README also says the software is "developed with strong assistance from AI coding agents". That matches the pattern in our coverage of Ponytail's 151,400-star surge, where agent tooling spread faster than anyone measured it. The repository showed 797 commits on main and 279 open issues.
Licences and weights
The engine is MIT. The pages fetched for this story do not state the licence of each supported model, and the README does not say where the quantized weights are hosted, so anyone planning commercial use has to check each model card. The Qwen3.8-27B release, for comparison, is a different, smaller Qwen model under Apache 2.0.
Stars are not adoption. The next thing to check is whether the 279 open issues, and the project's own beta label, change before anyone puts it behind a production API.
Sources
More in Software
- 01OpenCut Has 92,200 Stars, but the Editor People Use Is the Classic One and the Rewrite Is Not Taking ContributionsThe open-source CapCut alternative rebuilt its default branch in May. The README and a third-party walkthrough disagree on how much of the new code is Rust.
- 02Impeccable's Design Detector Runs Without a Model, but Its Open Issues Show Gaps Outside .htmlPaul Bakaus's design skill for coding agents ships 61 deterministic rules you can run from the command line. The bug tracker says where they are least reliable.
- 03A Hacker News Post Says Agents Need Documentation, Not Memory, and Its Author Wrote the Plugin That Does ThatKevin Liao's October 3 essay attacks snippet-recall memory plugins and promotes Operator Memory. A separate September essay argues the real gap is neither memory nor documents.
- 04Agent Reach, at 90,900 Stars, Reads X and Reddit for Your Agent Through Your Own CookiesThe MIT-licensed CLI routes agents to 20-plus sites with a backup backend per channel. Its README admits the login channels can get an account banned.