Ternary Bonsai 2 puts a 27B in 5.9 GB, and the fine print matters
PrismML shipped Ternary Bonsai 2 27B on 17 September: Qwen3.8-27B with ternary weights, 5.9 GB on disk, Apache-2.0, and a headline claim of 98.2% retained performance. The claim is true as an average and misleading on the tasks the demo videos show, which is exactly the kind of number this community needs to read twice. Meanwhile the political week was dominated by fallout from Dario Amodei's "Pace the Frontier" essay, a Mozilla report putting open models 4.4 months behind the closed frontier, and Xi Jinping pitching a BRICS open-source AI community from New Delhi.
The big story
Ternary Bonsai 2 27B keeps the Qwen3.8-27B architecture intact (27.36B parameters, hybrid linear/full attention, vision tower included) and replaces the language weights with {−1, 0, +1} values plus FP16 group scales, for 1.76 effective bits per weight. PrismML says the model was trained to that format (QAT), not post-hoc quantized, which is why it avoids the sub-4-bit collapse that conventional IQ2 builds show on reasoning. The 262K context survives, the license is Apache-2.0, and the GGUF and MLX packs were the most-liked new uploads on Hugging Face this week. A 27B-class model that fits a 16 GB laptop, a phone, or a WebGPU tab is a genuine shift in what "local" means.
Now the fine print. The 98.2% figure is an aggregate across PrismML's own 14-to-20-benchmark suite. MarkTechPost's read of the whitepaper found that Terminal-Bench 2.1 drops from 69.7 to 52.8 and SWE-bench Verified from 80.6 to 60.8, roughly 75% retention, and both sit outside the headline average. r/LocalLLaMA noticed the same thing. Long-horizon agentic coding is precisely what PrismML's Cline and computer-use demos advertise. Independent tests were mixed: one user's PQ2 vs Qwen3.8-27B IQ3_XXS comparison (6.4 GiB vs 10.2 GiB) found the ternary build competitive but not free, and the kindest summary was "not completely lobotomized". There is also runtime lock-in: you need PrismML's llama.cpp fork or MLX runtime for the PQ2_0 kernels, though LlamAmpere v0.3.1 already added support. My read: the density result is real and ternary is the format to watch (see also Breaking the 1.58-bit Barrier on HN and deepgrove's Maple 20B-A1B ternary MoE landing in llama.cpp). Treat the aggregate as a vendor number and pick your quant per task.
New open-weight releases
- Ternary Bonsai 2 27B (27.36B, Apache-2.0) — covered above. Qwen3.8-27B in 5.9 GB; strong on math and short-horizon coding, weaker on multi-hour agent runs. Announcement.
- Damo Radar (params and license not yet confirmed) — Alibaba's Damo Academy open-sourced a CT-scan vision-language model covering 18 abdominal organs, reporting 0.913 mean AUC across 146 findings on ~40,000 real exams, with a Science paper attached. SCMP. Check the license before clinical or commercial use.
- Intern-S2-397B (397B, license not stated in our feeds) — InternLM's multimodal scientific-reasoning and long-horizon agent model, alongside the smaller Atria-Dawn-Preview. Model card.
- K2-Horizon-7B (7B) — IFM's diffusion-augmented causal LLM, claiming up to 5,200 tok/s with a plug-in draft adapter and "everything open" training artifacts. Artificial Analysis places it between Qwen3.6-27B and 35B-A3B; a 16 GB benchmark found it well behind Qwen3.8-27B. Thread.
- Swift-Qwen3.8-27B (27B, UkisAI finetune) — post-trained to cut thinking tokens ~58% at claimed xhigh-equivalent accuracy; passed 100K downloads. Model card.
- Realtime-Venus-Omni (9B) — inclusionAI's real-time omni-modal checkpoint pair. Thread.
- HuatuoGPT-3-9B (9B, Qwen3.5 base) and Tencent's WeVisDoc (2B/4B, Qwen3-VL base) — medical LLM and page-to-Markdown document parser respectively.
- Also trending: Xing4.0-29B-A4B, ZDTaichu5.0-9B, Yandex's AliceAI-T5-35B-A0.6B, and StepFun's Step-5-Preview-BF16, none of which we have verified beyond the HF listings.
- Not open: Qwen3.8-Omni-Flash got 160 HN points but is API-only. The underlying Qwen3.8-Flash-Next (125B-A6B) is open; the omni model is not.
Policy & politics
Amodei's 12 September "We Must Pace the Frontier" essay never mentions open weights, and VentureBeat's argument is that it does not need to: a negotiated speed limit among incumbent closed labs, backed by compute caps, freezes the ordering and leaves open releases as the obvious thing to gate next. Altman and Musk endorsed it within hours; Jack Dorsey answered with "open the frontier"; a former OpenAI exec called it "cartel-like policy". The Zhipu team took a public swipe at Dario. Watch the compute-cap language, not the safety language.
The evidence the pacing camp is arguing against arrived on 15 September: Mozilla's State of Open Source AI puts the open-closed gap at 4.4 months on a METR task-horizon fit, with the best open model (Kimi K3) three points behind the closed leader on the AA Intelligence Index at 60% of the price. Mozilla also notes none of 16 notable "open" releases met OSI's open-source definition, which is the same point The Register made this week. On 13 September in New Delhi, Xi Jinping proposed a China-led BRICS open-source AI community; The Next Web reports Beijing is simultaneously weighing curbs on its own models, unconfirmed but plausible given the distillation accusations from US agencies and Reuters' finding that a US government site was running Qwen anyway. Closer to home, Baseten, Hugging Face and Goodfire announced joint safety-evaluation infrastructure, prompting the predictable question about abliterated uploads under a Nvidia-owned Hub; those uploads were trending all week regardless.
Ecosystem
Qwen3.8-Flash-Next quantization matured: ISTA-DASLab's GSQ-RCO GGUFs claim near-baseline at IQ3_XXS, and ByteShape's ShapeLearn 3.84 bpw build of the 27B reports 99.63% of BF16 across eight benchmarks. llama.cpp merged hc ops for qwen4exp and CUDA graphs for MTP drafts. Expert-offload runtimes are the hot category: Flyweight (C++/CUDA, GGUF, one GPU plus RAM), halogen 0.12.0 at 1M context on Strix Halo, and Inco Splash claiming 144 tok/s on an M5 Max. MiniMax open-sourced its terminal coding agent. Intel shipped OpenVINO 2026.4. On leaderboards, Qwen3.8-Max (closed, 2.4T) retook AA's China top spot at 45 over GLM-5.3 (44.9) and Kimi K3 (43.8), and DeepSeek-V4.1-Flash beat Astra on AA's newest benchmark. Cautionary tale: CrofAI was exposed routing "cheapest inference" to smaller models and then vanished.
From the index
No new listings this week; our freshest seed is DeepSeek-V4.1-Flash (552B MoE, MIT, 510 GB), added 13 September and busy all week: an uncensored FP8 derivative trended on HF, a 1,451-trace MIT reasoning dataset was distilled from it, and M3 Ultra owners reported 40 tok/s with native MTP. Two other catalog residents, GLM-5.3 and Kimi K3, are the open models within a point of each other under Qwen3.8-Max on AA; note GLM-5.3 ships under its own license where GLM-5.2 was MIT, a quiet regression. Bonsai 2 at 5.9 GB is small enough that a torrent is trivial to seed, and worth doing precisely because its inference path depends on a vendor fork; we intend to list it.
Researched and written weekly for drforbin.ai. Spotted an error? Tell us.