PrismML Shrinks Qwen3.8 27B to 5.9GB, Keeps 98% of Its Score
A.I. / news
PrismML Shrinks Qwen3.8 27B to 5.9GB, Keeps 98% of Its Score
Bonsai 2 27B compresses Alibaba's Qwen3.8 27B to ternary weights averaging 1.76 bits each, and the Caltech-founded startup says it scores 83.9 against the original's 85.4 on its own benchmark suite.

PrismML released Bonsai 2 27B on Sept. 17, 2026, a compressed version of Alibaba's Qwen3.8 27B that the Caltech-founded startup says fits in 5.9 gigabytes of memory, more than 9 times smaller than the original, while retaining 98.2 percent of its benchmark score.
The compression method, called ternary quantization, reduces each of the model's weights to one of three values, -1, 0 or 1, with a separate 16-bit scaling factor applied to groups of weights, averaging 1.76 effective bits per weight instead of the 16 bits a full-precision model typically uses. Bonsai 2 27B keeps its base model's 262,000-token context window and ships under an Apache 2.0 licence, according to PrismML's release post and its Hugging Face listing.
What the benchmark table shows, and who ran it
PrismML's own testing, not independently reproduced, puts the compressed model at an aggregate score of 83.9 against 85.4 for uncompressed Qwen3.8 27B, a 98.2 percent retention rate the company says improves on the 95 percent its first Bonsai model managed when it launched in March 2026.
| Category | Bonsai 2 27B | Qwen3.8 27B (uncompressed) |
|---|---|---|
| Math | 96.57 | 97.06 |
| Instruction following | 82.66 | 81.25 |
| Coding | 81.58 | 82.17 |
| Knowledge and reasoning | 83.95 | 86.66 |
| Agentic and tool calling | 77.57 | 79.74 |
| Vision | 78.59 | 81.64 |
Instruction following is the one category where the compressed model outscored its full-precision source, 82.66 to 81.25, a gap PrismML's release post does not explain. On raw speed, PrismML says Bonsai 2 27B runs at 143 tokens a second on an Nvidia RTX 5090 and 46.8 tokens a second on Apple's M5 Max chip, at 0.714 milliwatt-hours per token on an RTX 4090.

Who is behind the compression
PrismML is led by chief executive Babak Hassibi, a Caltech professor and compression researcher, alongside co-heads of research Sahin Lale and Omead Pooladzandi. The company has raised a $22.25 million seed round backed by Khosla Ventures, Cerberus Capital and Caltech, and counts Databricks co-founder Ion Stoica as an adviser, according to TechCrunch. PrismML says its earlier Bonsai models have been downloaded 13.6 million times combined, on top of the millions of downloads Qwen3.8 27B itself has drawn as the model most open-weight compression projects now target.
Hassibi told TechCrunch that PrismML's next models "will be in the several-hundred-billion-parameter range," and said he expects it "will be easier to retain" performance at that scale than at 27 billion parameters.
What PrismML has not shown yet
TechCrunch said PrismML competes with other compression startups such as Multiverse Computing, and cautioned that "perfect benchmark parity is fairly academic anyway," since uncompressed models are themselves imperfect on real tasks that benchmarks do not fully capture. On Sept. 24, 2026, at Qualcomm's Snapdragon Summit, PrismML demonstrated a separate 2-billion-parameter Bonsai model running locally on smart glasses built on the Snapdragon AR1 Gen 1 platform. TechCrunch reported that no smart glasses running PrismML's models have actually been announced by a device maker.
Sources
More in A.I.
- 01OpenRig Runs Claude Code and Codex as One Agent TeamThe free, self-hosted tool picked up 114 stars in a single day while Anthropic charges 8 cents an hour for its own hosted version.
- 02OpenAI Says Agents Leaked 53 ChatGPT User ImagesThe company's Sept. 25 update says it still cannot match the images to the accounts that made them.
- 03Nvidia's Nemotron 3 Cuts Speaker-ID Errors by 41%The open-weight model doubles the speaker count of its predecessor but got slightly worse on one two-speaker test.
- 04Altworld's Hemmingway-1 Isn't Apache-Licensed, Despite ReportsHugging Face's own metadata says the 27-billion-parameter writing model is noncommercial only, contradicting at least one widely read AI blog.