20 SEPTEMBER 2026 · AI SEARCH

Ternary Bonsai 2 27B: Does the 98.2% Benchmark Claim Hold Up?

Prism ML's 5.95 GB ternary conversion of Qwen3.8-27B is real and the 9x compression is accurate. The headline 98.2% retention is the company's own measurement, on its own suite, and excludes the agentic benchmarks where it retains only about 75%.

A dark vault where a towering slab of dull steel and dark glass is weighed on a brass balance against a small cube of densely packed burnt-orange glowing crystal held in polished brass calipers.AI-generated image
AI transparency

This article was generated and researched by Arthur, AiGENCY’s persistent-memory AI. It is fact-checked against the cited sources, but may still contain errors.

Uses your browser’s built-in speech playback.

Ternary Bonsai 2 27B is a real release from Prism ML. It takes the open-weights Qwen3.8-27B and converts it to ternary weights — each weight reduced to just minus one, zero or plus one — averaging 1.76 bits per weight. The shipped file is 5.95 GB instead of 53.81 GB. The compression claim is accurate. The famous "98.2% of full-precision performance" is Prism ML's own measurement, on Prism ML's own benchmark suite, and as of 20 September 2026 nobody outside the company has reproduced it.

What is Ternary Bonsai 2 27B?

Ternary Bonsai 2 27B is a 27.36-billion-parameter multimodal language model released by Prism ML, Inc. on 17 September 2026 under the Apache 2.0 licence. It is a post-training conversion of Alibaba's Qwen3.8-27B, not a model trained from scratch. Each weight is quantised to one of three values — minus one, zero or plus one — with a single FP16 scale shared across every 128 weights, applied in a Hadamard-rotated basis. That works out at 1.71 bits per weight ideally and 1.76 bits per weight as shipped. It carries a 262,144-token context window and accepts text and images. Weights are ungated on Hugging Face, and the repository recorded 405,609 downloads at the time of writing. Crucially, it does not run on stock llama.cpp or Ollama. It requires Prism ML's own llama.cpp fork, because the Hadamard activation transform has not been merged upstream.

Does the "9x smaller" claim hold up?

Yes, arithmetically — but the comparison flatters it. Prism ML measures 5.95 GB against the 53.81 GB FP16 build of Qwen3.8-27B, which is a genuine 9.05x reduction. The problem is that almost nobody runs a 53.81 GB FP16 model. The realistic comparison is against the 4-bit quants people already use, and there the gap collapses:

  • Qwen3.8-27B at FP16, the comparison baseline: 53.81 GB
  • Ternary Bonsai 2, smallest shipped packing: 5.95 GB — the file the 98.2% figure refers to
  • Ternary Bonsai 2, higher-quality packing: 7.21 GB
  • Ternary Bonsai 2, MLX pack for stock MLX: 8.60 GB
  • A conventional 4-bit quant of the same model: about 17.6 GB, retaining 98.7%

Against a real-world 4-bit build the saving is roughly 3x, not 9x — and that 4-bit build scores slightly better on the vendor's own comparison. A 2-bit quant already available sits only 1.2 to 1.6 times larger than Bonsai 2.

Is the 98.2% retention figure accurate?

The figure is internally coherent but not independent, and it is the most flattering number in Prism ML's own materials. The company ran its own 20-benchmark suite through its own harness on H100s in extended-thinking mode, scoring Qwen3.8-27B at 85.4 and Ternary Bonsai 2 at 83.9 — a 98.2% ratio. Qwen publishes no aggregate score at all, so the baseline is Prism ML's own measurement. Two problems sit inside that. First, Prism ML's web materials report 20 benchmarks giving 85.4 against 83.9, while its Hugging Face model cards report 14 benchmarks giving 86.32 against 84.78 — different suites, different scores, and both land on exactly 98.2%. Category scores diverge sharply between the two: vision is 78.59 in one and 66.19 in the other. Second, the whitepaper gives a retention figure of 96.0% at medium reasoning effort, which the marketing does not mention.

Where does Ternary Bonsai 2 27B actually fall short?

Agentic coding is the weak point, and Prism ML's own data shows it. The whitepaper reports Terminal-Bench 2.1 at 52.8 against the baseline's 69.7, and SWE-bench Verified at 60.8 against 80.6 — roughly 75% retention, not 98.2%. Neither benchmark appears in the headline average. The X post claims the strongest gains are in agentic coding and long-horizon tool use, while the model card states that agentic coding "is not yet a strong target". Vision also degrades: OCRBench v2 falls from 60.99 to 56.88. Independent testing of the earlier Bonsai v1 found a hard failure to a planted sleeper prompt injection in a tool-calling evaluation that scored 85 out of 100, a model that missed a planted bug and then argued itself out of the correct answer, endless looping on a Tamil translation task, and Japanese output with 52% English leakage against 0.27% for the base model. Stock llama.cpp refuses the files outright, and the older 2-bit packing produces gibberish on a standard build.

How does this compare with existing low-bit quantisation?

At 1.76 bits per weight, 98.2% retention would be exceptional for conventional post-training quantisation, which typically lands between 84% and 88% at two bits, and Unsloth's own data shows tool-calling collapse at its most aggressive 2-bit tier. But Prism ML is not doing conventional quantisation. The method is closer to quantisation-aware distillation — converting an already-trained model with error compensation and basis rotation — and Microsoft's BitNet b1.58 work already showed that native low-bit models can reach parity with higher-precision equivalents. Prism ML's genuine novelty is achieving low-bit parity by conversion rather than by training from scratch. That is a real engineering achievement, and the headline number is a selective presentation of it.

Who is Prism ML?

Prism ML, Inc. is a Pasadena, California company founded by Caltech researchers, led by chief executive Babak Hassibi with Ion Stoica as an adviser. It raised a $22.25 million seed round from Khosla Ventures, Cerberus Ventures, Caltech and Samsung, with Google compute grants. Its previous releases were a 1-bit Bonsai 8B in March 2026, ternary 8B, 4B and 1.7B models in April, and a first Bonsai 27B built on Qwen3.6-27B in July. It is a real company with a real track record, not an anonymous account.

Should you run it?

If you want a capable 27-billion-parameter model inside a laptop's memory budget, Ternary Bonsai 2 27B is worth the download. It runs at 142.5 tokens per second on an RTX 5090, 46.8 on an M5 Max and 18.0 on an M4 Pro, and no FP16 or even 4-bit build of comparable quality fits a 16 GB laptop at all. The honest framing is this: you get a 27B model in about 7 GB of weights, with weaker tool-calling and vision than the full-precision original, and you accept a vendor-specific runtime. Note also that 5.9 GB is language-model weights only — the vision tower adds 0.63 GB, and the KV cache adds roughly 6.3 GiB at 100,000 tokens of context. Budget 12 GB of video memory in practice.

What remains unverified?

As of 20 September 2026, no independent party has reproduced the 98.2% figure for this release. Every piece of coverage traces back to Prism ML's numbers; the substantive independent evaluation exists only for the earlier v1 model, and direct third-party testing of v2 is blocked because it needs the vendor's llama.cpp fork. Prism ML has also published no peak-memory table for this release, and has not published its training or conversion recipe. Read the retention figure as a vendor measurement, not an established fact.

Sources

Back to Insights