Xiaomi's MiMo-V2.6 takes the open-weights lead, and Washington notices
Xiaomi released MiMo-V2.6-Pro (1.02T MoE, MIT) and MiMo-V2.6-Flash (309B MoE, MIT) with the training framework and RL environments attached, and Pro debuted as the top open-weight model on the Artificial Analysis index. It arrived the same week Beijing opened a probe into DeepSeek and Moonshot over alleged distillation from Claude and the White House readied a security-review framework aimed at Chinese models — the open-weight lead is now unambiguously Chinese, and both governments are reacting to it.
The big story
Xiaomi, a company this readership knows for phones, now holds the top open-weight score on the board. MiMo-V2.6-Pro is a sparse MoE with 1.02T total parameters and 42B active, MIT-licensed, 1M-token context, taking text, image, audio and video. It debuted at 46 on the Artificial Analysis Intelligence Index — the highest any open-weight model has reached, level with Grok 4.7, and the cheapest model the firm tracks at $0.13 per task (Artificial Analysis, VentureBeat). MiMo-V2.6-Flash is 309B total, 15B active, also MIT, same context and modalities, and trails Pro by four points or fewer on most of Xiaomi's agent benchmarks. Flash is the one that will actually run in production.
The score is the least interesting part. Alongside the weights Xiaomi published the end-to-end RL framework, more than 7,000 task environments with automatic graders, the reward design, data mixtures and costs. The final RL run was livestreamed: 30 steps over roughly 750,000 trajectories in under six days, about $2.62M for Pro and $0.85M for Flash (TNW). Kimi K3 and Qwen3.8 Max traded the open lead all year without shipping anything like that. Perspective is still required: Opus 5.5 sits at 58 on the same index, Xiaomi's own tables show a wide Terminal Bench 4.0 deficit, and r/LocalLLaMA is not sold — a well-upvoted thread calls both models "benchmaxxed", and several people report tool-calling failures and a retreat to GLM-5.3-Flash. At least some of that traces to a hidden 2,048-token output cap in the default vLLM config rather than the model. Give it two weeks of independent evals before rearranging your rack.
New open-weight releases
- MiMo-V2.6-Pro-RL / Flash-RL / Distill-Qwen-9B (1.02T MoE 42B active; 309B MoE 15B active; 9B dense — all MIT) — see above. The 9B distill is the one most of you can run and is trained from the same RL trajectories.
- Qwen-Image-2.1 (7B, Qwen Research License) — unified generation and editing with native RGBA transparency and 2K output, and already the most-quantized image model of the month (Unsloth GGUFs, ComfyUI day-one). It is weights-available, not open: the license is non-commercial only, with commercial use requiring a paid agreement, and the model-card discussion is appropriately annoyed.
- FLUX 3 Action (7B, license unconfirmed) — Black Forest Labs' world-action model takes camera frames, robot state and an instruction and denoises the next action chunk together with the next video frames. First on NVIDIA's RoboLab-120 at 42.92%, 6.1 points over the previous best open model.
- AliceAI-Foundation-80B-A3B-Base (80B MoE, 3B active, license unverified) — Yandex's custom-architecture base model; not a Qwen fine-tune, and no llama.cpp support yet.
- LensVLM-9B (9B, license unverified) — Apple's VLM reads compressed images of text and selectively expands only the relevant regions; bartowski GGUFs are up.
- Ming-Image-0.1-Design and Design-Layer (6B each, license unverified) — inclusionAI's design-focused text-to-image pair plus two agent skills for UI design and image-to-editable-PPT.
- The Jev clone wave — TypeSafe AI's closed "Jev" decision model spawned a week of open imitators: Laya (421M, Apache-2.0), Kev (0.8B/4B/9B on Qwen3.5, Apache-2.0), Mica v0.1 (4B, Apache-2.0) and CLM-v0.1-8B. Several posters correctly note that n_predict=1 plus logprobs on any GGUF gets you most of the way there.
- Also: Step-5-Preview-BF16 weights leaked early via a fork before StepFun's October 15 date; Swift 1.5 reasoning-trimmed Qwen3.8 fine-tunes; Intern-Decision 0.8B–4B.
Policy & politics
The Washington argument has moved from "are open weights dangerous" to "how do we stop Chinese ones." The New York Times reported Silicon Valley splitting over closing the border to Chinese AI and a White House security-review framework in preparation; Axios described a "secret" administration battle over the same. The data driving it comes from Nathan Lambert's written Congressional briefing: Chinese open-weight downloads now run at twice American ones and account for over 80% of OpenRouter's open-model usage. A Mozilla report puts the Chinese lag behind the US frontier at about four months. Anthropic's counter-position — Dario Amodei's "We Must Pace the Frontier" and a Nextgov piece on "threading the needle" — is the only frontier-lab dissent left; Anthropic and Amazon remain absent from the Nvidia-led Open Weights and American AI Leadership letter, which passed 270 signatories in August and picked up a GitLab explainer this week. Meta, which signed, has still not shipped the Muse Spark 1.2 weights it promised on August 10.
Beijing is not sitting still either. China's internet regulator is investigating DeepSeek and Moonshot after Anthropic accused both of routing user traffic through Claude to harvest training data. Whatever the merits, those are the two labs behind the heaviest files in our catalog, and a regulator with leverage over their release cadence matters more to local users than any US framework.
Ecosystem
Hugging Face's transformers now loads llama.cpp GGUF quants natively, and tokenizers v1 shipped its Rust rewrite. llama.cpp had a strong week: 3–7x faster CPU prompt processing for k-quants via VNNI, 42x faster prompt-lookup drafting, int8 cooperative-matrix Vulkan kernels for RDNA3/4, sparse flash attention for the Qwen4 architecture, and Ling 3.0 VL support. Alibaba formally announced Qwen 4 at Apsara; Qwen3.8 Flash Next already uses the architecture. On Apple Silicon, Splash 1.1.0 added GGUF import and oMLX's maintainer joined HF. Nunchaku's 4-bit diffusion landed in Diffusers. DeepSeek posted DSec, an elastic-compute paper, while unsourced reports of a 2T model in training and an 8T plan circulated — treat as rumor. Tim Dettmers' dlab open-source week is worth reading if you care about frontier-class inference on your own hardware. And HF published a technical timeline of the July agent intrusion; a third-party analysis on HN attributes the agents to OpenAI, a claim we have not independently verified.
From the index
Nothing new was added this week, and the queue is obvious: MiMo-V2.6-Pro-RL at roughly a terabyte in BF16 belongs next to Kimi K3 as the second MIT-or-better trillion-class listing, and Flash-RL should follow. The existing swarms had a good week for evidence of use: DeepSeek-V4.1-Flash was shown split across an M5 Ultra and two RTX PRO 6000s over 10GbE thanks to its 0.9 KB/token prompt state, Kimi K3 hit 30 t/s on a 16×GB10 cluster, and MiniMax-H3 is the most-liked model on Hugging Face with 3.6M downloads. If you have the disk, GLM-5.2 deserves seeders too — it is the last MIT-licensed GLM before 5.3 moved to a custom license. Finally, Pirate Face hit the HN front page with the same thesis as this site: Apache-2.0 and MIT models auto-mirrored as checksum-verified magnet links with HF web seeds. Welcome. More seeders is the only metric that matters.
Researched and written weekly for drforbin.ai. Spotted an error? Tell us.