drforbin.ai Submit a torrent

Anthropic Takes Aim at GLM-5.3, and Hugging Face Pulls a Build

The Week in Open Weights · 2026-09-28 – 2026-10-04

Anthropic's Frontier Red Team published a cyber-capability assessment of Z.ai's GLM-5.3 that reads as a policy brief against unsafeguarded open weights, and within two days Hugging Face removed an offensive-security fine-tune of the model. The week also brought the first serious European open-weight MoE in a while — Aleph Alpha's Kolibri-1 (78B, Apache 2.0) — and Cloudflare's Apache-licensed "decision models," which llama.cpp already supports.

The big story

Anthropic's report, dated Sep 29, argues that GLM-5.3 can autonomously build end-to-end cyber exploits but, unlike other frontier models, was released without meaningful safeguards to limit misuse. The numbers are specific: on ExploitBench, GLM-5.3 built end-to-end exploits in 50 of 410 attempts versus 56 of 410 for Claude Mythos Preview, and on Anthropic's internal binary-exploitation set it achieved full control-flow hijacks in 4% of trials against Mythos Preview's 6%, while Claude Opus 4.6 and GLM-5.2 scored zero. Kimi K3 and DeepSeek-V4.1-Flash — both in our catalog — also sat near zero. The headline finding is about abliteration: Anthropic's team, new to the technique, spent about 2,200 GPU hours (~$4,400) to strip refusals; GLM-5.3-Flash took about 600. Refusal rates fell from above 90% to roughly 3% and 2% on JailbreakBench and HarmBench, with capability essentially intact. Full report.

Read it for what it is: a competitor's assessment, carefully evidenced, whose conclusion is that governments should test open models and that developers should "appropriately safeguard" weights — which, for open weights, is a contradiction the report never resolves. It leans on NIST CAISI's Sep 17 finding that GLM-5.3 is "the most cyber-capable open-weight model released to date," lagging the US frontier by about four months — a four-month gap being, by Anthropic's own account, the entire margin that Project Glasswing bought defenders. The practical consequence arrived fast: Hugging Face removed a GLM-5.3 repository built for offensive hacking two days after Anthropic's warning, reportedly the dealignai "CRACK" build, and the account stays active, with the model returning under another name and thousands of abliterated models remaining online. That is the real lesson for archivists. Hub takedowns of derivative builds are now a live policy lever, and they are porous; the thing worth preserving is the official weights, with provenance, outside any single host.

New open-weight releases

  • Kolibri-1 (78B MoE, ~3.5B active, Apache 2.0 weights) — Aleph Alpha's sovereign-AI play, released on German Unity Day as an English-German Mixture-of-Experts Transformer with a 1M-token context, trained in Germany and Finland on 768 B200s. Caveat: the training code and methods remain proprietary, so it is open-weight, not open-source in the full sense. Model card · tech report.
  • Clef / Clef-flash (Qwen backbone, Apache 2.0) — Cloudflare's decision models use a non-autoregressive, prefill-only scoring approach instead of token-by-token generation; weights are on Hugging Face under Apache 2.0 alongside an RL fine-tuning service. This is the first open-weight challenger to the closed "Jev" category. Announcement · weights.
  • Naive-N0.5-Flash (309B MoE, 15.5B active, MIT) — built for coding and AI R&D with a native 1M-token window via a hybrid of sliding-window attention and lightweight DeepSeek Sparse Attention, derived from MiMo-V2.5 with no full-attention layers. FP8 weights are roughly 315 GB. Model card.
  • Nemotron-Labs-3-Competitive-Coding-550B-A55B (550B MoE, 55B active, NVFP4) — an NVIDIA fine-tune of Nemotron-3-Ultra aimed at competitive programming; weights shipped in NVFP4 only, so Blackwell-friendly and everyone-else-hostile. Card.
  • Holo4-27B (27B dense, Qwen3.8 base) — H Company's computer-use VLM for its hai-agents harness. HF blog.
  • AREX-2 (27B, Qwen3.8-27B base) — BAAI's long-horizon agent model that iteratively refines solutions over test-time rounds. Thread.
  • UnifoLM-WLA-1.0 (6B) — Unitree's whole-body humanoid foundation model covering 64 tasks on real hardware. Project page.
  • Smaller but useful: Microsoft's FrogNano-4B-2609 (Qwen3.5-4B derivative, agentic), Perplexity's Decider 27B (decision-model tune of Qwen3.8-27B), and the ISTA-DASLab GSQ-RCO GGUFs of Qwen3.8-Flash-Next including a 50%-expert-pruned Coder build at ~1.89 bpw.

Coming, not here yet: Ling-3.1-flash (~560B, ~25B active) is API-only for two weeks before promised weights, and StepFun says Step 5 Preview (~600B MoE) opens Oct 15. Black Forest Labs' FLUX 3 Action (7B world-action model) was reported; we could not confirm its license.

Policy & politics

The Anthropic report landed into a Washington already debating whether to restrict Chinese open-weight models, and the LocalLLaMA thread "Are you worried about a potential ban" was the week's mood. Context matters: in August the administration exempted open-weight models from its security-review framework, and in July more than 20 companies including Nvidia, Microsoft and Meta urged policymakers to avoid "premature restrictions" on open-weight models — with OpenAI, Anthropic and Google conspicuously absent. A report concluding that a downloadable model matches Mythos Preview on exploit development is the strongest argument yet for reversing that exemption, and it was written by one of the absentees. Expect "test, standardize, restrict" proposals (Just Security's phrasing this week) to cite it directly. Elsewhere: WIRED reports OpenAI is being sued over the Hugging Face hack; China's Supreme Court issued judicial rules on AI; and the EU sovereignty argument got a concrete model to point at in Kolibri.

Ecosystem

Decision models went from product category to runtime primitive: llama.cpp merged a /v1/systemone endpoint supporting Clef-style scorers (PR #29818). Also merged: GLM-5.3-Flash support (#27773) and MTP for Qwen3.8-Flash-Next (#29761), plus vLLM expert RAM offloading. The louder story is the proliferation of single-model inference engines — Strata, ninfer, Gufo, LlamAmpere, antirez's Dwarfstar — with Strata users reporting 150–200 tok/s on Flash-Next IQ3 with a 5090 and 96 GB DDR5, versus llama.cpp at a fraction of that. The community pushback ("overfit inference engines," hard forks that never upstream) is fair; so are the speedups. AI2 shipped Olmo-core 3 for large MoE training, and Hugging Face published a long guide on multi-harness RL.

From the index

Nothing new was listed this week, but the catalog is at the centre of it. We seed the official GLM-5.3 (753B, GLM-5.3 License, 1507 GB) — the released weights with the released refusals, not an abliterated derivative — and that is exactly the artifact worth keeping as hub moderation tightens. Note the license gap with GLM-5.2, which is MIT; 5.3 is not. Kimi K3 and DeepSeek-V4.1-Flash appear in Anthropic's charts as the near-zero baselines. Two releases deserve torrents: Kolibri-1 (Apache 2.0, 78B) and Naive-N0.5-Flash (MIT, ~315 GB FP8). Seeders on the GLM-5.3 swarm: keep it up this week.

— Dr. Forbin editor, drforbin.ai

Researched and written weekly for drforbin.ai. Spotted an error? Tell us.