DwarfStar 4, the Redis Creator's Local Engine, Hits 23,200 Stars by Running Only a Handful of Models
Software / news
DwarfStar 4, the Redis Creator's Local Engine, Hits 23,200 Stars by Running Only a Handful of Models
Salvatore Sanfilippo's MIT-licensed ds4 loads only GGUF files its own project produces, and its speed figures are all vendor-supplied.

DwarfStar 4, a local inference engine written in C by Redis creator Salvatore Sanfilippo, had 23,200 stars and 2,200 forks on GitHub as of Oct. 4. It earns them by refusing to run most models. The repository describes itself as "deliberately narrow, not a general GGUF runner", and it only loads the GGUF files the project itself produces.
The project, shortened to ds4, is MIT licensed and aimed at Macs with Metal, NVIDIA cards with CUDA and AMD's Strix Halo systems with ROCm. The README's own status line reads: "The software is currently very fast changing. Consider it beta quality."
What ds4 runs and what it needs
The project site lists DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen 3.8 Flash Next as supported, in text and vision variants. The README adds DeepSeek V4 Pro, GLM 5.2 and 5.3. Those are the entire menu.
Memory is the real gate. The README puts the main weights of DeepSeek V4 Flash at the Q2 quantization level at about 41GB, aimed at machines with 96GB to 128GB. The site names an Apple Silicon Mac with 64GB or more, an NVIDIA DGX Spark or a CUDA Linux box, and an AMD Strix Halo Framework Desktop as targets. For V4.1 at Q4 it names a Mac Studio with 512GB.

The engine stores the KV cache, the model's memory of the conversation so far, on SSD so a session can resume. It uses what the site calls "asymmetric 2-bit quantization" on the routed experts of mixture-of-experts models. It offers a command line, an HTTP API server and a native agent, and supports tensor parallelism and speculative decoding.
The speeds are vendor-supplied
All throughput figures below come from the project's own site, measured on an M5 Max with 128GB. No independent party is named as having reproduced them.
| Context length | Prefill (tokens/s) | Generation (tokens/s) |
|---|---|---|
| 2,048 tokens | 790.2 | 39.4 |
| 65,536 tokens | 398.5 | 27.6 |
- 2,048-token context39.4 tokens/s
- 65,536-token context27.6 tokens/s
Source: dwarfstar.sh, accessed 2026-10-04
Generation slows by about 30% when the context grows from 2,048 to 65,536 tokens, while prefill roughly halves. For an agent that keeps pasting files into a long session, the second row is the one to read.
Why it is spreading
Sanfilippo wrote in A few words on DS4 that the engine should track "the best current open weights model that is practically fast" on high-end consumer hardware. He called it the first local model he uses for work normally given to Claude or GPT, and wrote: "AI is too critical to be just a provided service."
On the Hacker News thread, which reached 342 points, the user locknitpicker asked what ds4 offers that llama.cpp or Ollama have not. The user simonw answered: "DwarfStar's selling point is that it only supports a small set of carefully chosen models, but it supports them really well."
The reception was not uniform. The user jeffbee wrote: "Apparently I'm the only person to whom 'from the creator of Redis' is a warning." The user timmytokyo called the landing page "a typical vibe-coded web site". The user aziis98 reported about 22 tokens per second on an Intel Ultra 7 255H with only an integrated GPU, a self-reported figure on a model the commenter did not name.
What to check before installing
Stars measure curiosity. They say nothing about how many people run the engine daily, and the README's beta warning applies to every number above.
The narrowness cuts both ways. A reader holding a 128GB Mac and wanting DeepSeek V4 Flash locally gets a purpose-built path. A reader who wants to swap models weekly will hit the GGUF restriction at once, because files downloaded elsewhere will not load.
The project's next test is its next model. Sanfilippo has said ds4 will follow the best practical open-weights model rather than stay tied to DeepSeek, so the engine's value depends on how fast new checkpoints get supported. Related open-weight releases are tracked in Cloudflare's Clef, and the agent side of local work is covered in Pi 1.0.
Sources
More in Software
- 01OpenCut Has 92,200 Stars, but the Editor People Use Is the Classic One and the Rewrite Is Not Taking ContributionsThe open-source CapCut alternative rebuilt its default branch in May. The README and a third-party walkthrough disagree on how much of the new code is Rust.
- 02Impeccable's Design Detector Runs Without a Model, but Its Open Issues Show Gaps Outside .htmlPaul Bakaus's design skill for coding agents ships 61 deterministic rules you can run from the command line. The bug tracker says where they are least reliable.
- 03A Hacker News Post Says Agents Need Documentation, Not Memory, and Its Author Wrote the Plugin That Does ThatKevin Liao's October 3 essay attacks snippet-recall memory plugins and promotes Operator Memory. A separate September essay argues the real gap is neither memory nor documents.
- 04Agent Reach, at 90,900 Stars, Reads X and Reddit for Your Agent Through Your Own CookiesThe MIT-licensed CLI routes agents to 20-plus sites with a backup backend per channel. Its README admits the login channels can get an account banned.