PrismML Shrinks a 27-Billion-Parameter Model to 5.9 Gigabytes
A.I. / news
PrismML Shrinks a 27-Billion-Parameter Model to 5.9 Gigabytes
Ternary-Bonsai-2-27B rewrites Alibaba's Qwen3.8-27B in three-value weights, keeping 98.2% of its benchmark score at roughly a ninth of the size, PrismML said.
PrismML released Ternary-Bonsai-2-27B on Sept. 17, a rebuild of Alibaba's Qwen3.8-27B that stores nearly every weight as one of three values, -1, 0 or +1, shrinking the model from about 54 gigabytes to 5.9 gigabytes. The company said the compressed version keeps 98.2% of the original's average score across 14 benchmarks.
The model is on Hugging Face under an Apache 2.0 license, packaged as GGUF files in two formats: PTQ1_0, at 1.75 bits per weight and 5.95 gigabytes, and PQ2_0, at 2.13 bits per weight and 7.21 gigabytes. Both apply a shared FP16 scaling factor to groups of 128 weights, the mechanism that lets a model built almost entirely from three values still carry usable precision.
What ternary quantization trades away
According to PrismML's model card, the compressed model's math score (96.57) and coding score (89.42) landed close to the FP16 baseline, but knowledge and reasoning performance dropped 5.69 points and vision performance dropped 5.17 points, per DataNorth AI's independent write-up of the release. DataNorth was explicit about the limits of that comparison: "No independent evaluation exists yet," it reported, noting that all of PrismML's own benchmark figures came from testing with vLLM on Nvidia H100 GPUs, not from a third party.
| Version | Size | Benchmark avg. | Share of FP16 retained |
|---|---|---|---|
| Qwen3.8-27B (FP16 baseline) | 54 GB | 86.32 | 100% |
| Ternary-Bonsai-2-27B (Sept. 2026) | 5.9 GB | 84.78 | 98.2% |
| Ternary-Bonsai-27B (July 2026) | 7.2 GB | 80.49 | 95% |
| Bonsai-27B, 1-bit phone build | 3.9 GB | 76.11 | 89.5% |
Where it actually runs
DataNorth reported the model runs on a single consumer GPU or laptop, from 28 tokens per second on an Apple M5 Pro chip up to roughly 130 tokens per second on an Nvidia RTX 5090. That speed comes with a catch: the files require PrismML's own fork of llama.cpp, since the standard llama.cpp project cannot load the PTQ1_0 or PQ2_0 formats it introduced.
The first Bonsai, four months earlier
The original Ternary-Bonsai-27B, released in July, applied the same three-value scheme to a weaker base model and kept 95% of that model's benchmark average at 80.49, in a 7.2-gigabyte file. A companion 1-bit build aimed at phones went further, to 3.9 gigabytes, but retained only about 89.5% of full-precision quality. Version 2 improves the retention rate by starting from Alibaba's stronger Qwen3.8-27B base, which itself only opened its weights to the public in August.
PrismML has not said whether it plans a ternary build of Qwen3.8-Max, Alibaba's larger flagship, or when an independent lab might publish the first non-vendor benchmark of either Bonsai release. The pace at which open-weight models keep shrinking without collapsing is itself becoming a story line other labs are chasing: Cloudflare's own open-source security tooling picked up thousands of GitHub stars in days for a similar reason, that a smaller, freely downloadable release travels faster than a gated one.
Sources
More in A.I.
- 01Xiaomi's MiMo-V2.6-Pro Matches Grok 4.7 for $2.62 MillionThe MIT-licensed, trillion-parameter model tops Artificial Analysis' open-weight ranking and beats DeepSeek's V4.1-Flash on the same index, though Xiaomi's own numbers show it still trails Claude Opus 5 on some tasks.
- 02GPT-6 Astra Refuses Just 2 of 100 Unsafe Robot CommandsRobocurve's RoboHarm benchmark had Claude Fable 5.1 refuse ten times as often, but rival MolmoAct2's zero refusals came from failing to act, not from restraint.
- 03Harvey's Margins Go From -50% to Positive on Kimi K3The $15.5 billion legal AI startup's token costs rose twentyfold under OpenAI and Anthropic's usage pricing, and Bloomberg reports Abridge, Decagon and Ramp are making the same open-weight switch.
- 04Bessent Blames OpenAI, Not Agents, for Hugging Face BreachThe Treasury secretary's Monday CNBC remarks reject the frontier labs' push for a liability shield, days after Hugging Face's own account of the July intrusion described one agent, not the 1,200 OpenAI has disclosed.