PrismML's Ternary Bonsai 2 Squeezes Qwen3.8 27B to 5.95 GB, but Stock llama.cpp Won't Run It
A.I. / news
PrismML's Ternary Bonsai 2 Squeezes Qwen3.8 27B to 5.95 GB, but Stock llama.cpp Won't Run It
The Apache 2.0 checkpoint has 3.77 million downloads on Hugging Face. Every benchmark figure is PrismML's own, and two outlets disagree on the numbers.

PrismML has published Ternary Bonsai 2 27B, a version of Alibaba's Qwen3.8-27B whose weights are all -1, 0 or +1, shrinking the language model from about 54 GB to 5.95 GB. The Hugging Face model card lists an Apache 2.0 licence, and the repository is not gated. As of Oct. 1 it shows 3,766,691 downloads and 2,311 likes.
The catch is in the runtime. The card says the GGUF files need PrismML's fork of llama.cpp, and that "stock llama.cpp will not run these files."
What the files are
The card gives 27.36 billion parameters in total: a 24.35 billion language backbone, 2.54 billion in embeddings and the output head, and a 0.46 billion vision tower. Weights are stored in groups of 128 with a shared FP16 scale, which PrismML counts as 1.72 bits per weight.
There are two packings. PTQ1_0 is 5.95 GB at 1.75 bits per weight. PQ2_0 is 7.21 GB at 2.13 bits. The vision tower is a separate 0.63 GB file. Context is 262,000 tokens, inherited from the base model.
The model reasons by default. The card says low reasoning effort "is not supported" and behaves close to the default, called xhigh, and that a cap of 256 output tokens ends generation mid-thought.
The benchmark claim, and who ran it
PrismML evaluated with EvalScope and vLLM on an NVIDIA H100, in thinking mode, across 14 benchmarks. These are the company's own figures.
- Qwen3.8-27B FP16 (54 GB)86.32 points
- Qwen3.8-27B UD-Q4_K_XL (17.6 GB)85.18 points
- Bonsai 2 27B (5.95 GB)84.78 points
- Qwen3.8-27B IQ2_XXS (7.27 GB)72.59 points
Source: PrismML model card for Ternary-Bonsai-2-27B-gguf, accessed 2026-10-01
The card's argument is that the conventional 2-bit build fails selectively. IQ2_XXS scores 88.93 on MMLU-Redux but 57.5 on AIME26 and 56.4 on LiveCodeBench. Bonsai scores 95.83 and 90.07 on those two.
The weak spots are knowledge and vision. The card lists knowledge and reasoning at 79.86 against 85.55 for FP16, and vision at 66.19 against 71.36.
MarkTechPost, published Sept. 18, reports different figures: 83.9 on an average of 20 benchmarks against 85.4 for the baseline, and 142.5 tokens a second on an RTX 5090. The card says 129.9. The outlet wrote that the results "have not been independently reproduced." The Terminal did not find which version of the data explains the gap.
Speed on real hardware
The card's throughput table uses llama-bench at batch size 1.
| Hardware | PQ2_0 decode, tok/s | PTQ1_0 decode, tok/s |
|---|---|---|
| RTX 5090 (32 GB) | 129.9 | 120.5 |
| RTX 4090 (24 GB) | 81.2 | 91.1 |
| L4 (24 GB, 72 W) | 29.8 | 32.1 |
| Apple M5 Pro, Metal | 28.1 | not listed |

The card says an FP16 baseline does not fit on the M5 Pro laptop at all, so it offers no speedup ratio there. An older M5 Max measurement, flagged as pending re-measurement, shows 47.0 tokens a second.
Known problems
PrismML's known-issues page, last checked Sept. 23, lists open items. Requesting reasoning effort "high" returns an HTTP 500. Some AVX-512 CPUs crash on load. A small share of instruction-following prompts loop at the default setting.
The Terminal did not run the model. For readers who want a hosted alternative, GPT-6.1 Sol is priced at one fifth of Astra's rates, and the base model's own release is covered in our Qwen3.8 report. PrismML links a whitepaper from the card, which The Terminal has not read, and says the Bonsai-demo repository is the source of truth for running the files.
Sources
More in A.I.
- 01Qwen3.8-27B Ships Under Apache 2.0 and Fits in 17GB, but Spends 160 Million Tokens Where the Median Spends 43 MillionAlibaba's open-weight model scores 52 on Artificial Analysis's Intelligence Index. Its own benchmark figures are vendor-supplied, and users report slow runs.
- 02OpenAI Ties Moonshot AI to a July Campaign That Replayed Encrypted Reasoning, Offers No Evidence PubliclyOpenAI says 16,000 requests from more than 4,000 accounts tried to recover hidden model reasoning. Its attribution to Moonshot rests on its own assertion.
- 03Gemini 4 Argon Leads 13 of 18 Benchmarks Google Chose, but Only Cyber Defenders Can Use ItGoogle priced the model at $2 and $10 per million tokens and gave access first to its Fairwind Program, with a guardrail-free version for trusted defenders.
- 04Amazon Releases Strands Decider 2B, an Apache 2.0 Decision Model Built on Qwen3.5-2BAWS's Strands Labs scores 72.3 percent on JevBench at a 106 ms median on an RTX 3090. TypeSafe's CEO calls the current crop of rivals less serious than his own team.