The Hub Has an Owner Now — Nvidia Pays $12.93B for Hugging Face
Nvidia confirmed on Thursday that it is acquiring Hugging Face for $12,930,300,000, putting the de facto distribution layer for open-weight models inside the company that sells the hardware they run on. The same week produced an unusually strong release slate — a 770B Apache-2.0 MoE from Tencent, DeepSeek's first V4 vision model under MIT, and a six-model fully-open fleet from the UAE's IFM — which is exactly the kind of catalog that now depends on one corporate owner's goodwill to stay reachable.
The big story
Nvidia's announcement, written by Jensen Huang himself, is careful where it needs to be: Hugging Face keeps its brand, stays multi-cloud and multi-accelerator, and "NVIDIA compute will not be required to build on or deploy through Hugging Face." The price — $12,930,300,000 — encodes 129303, the decimal value of U+1F917, the 🤗 emoji, which tells you how much the deal was designed as a community-relations exercise. Per TechCrunch, Hugging Face turned down a $500M Nvidia offer last year and was running roughly $150M in annualised revenue on 18 million developers, 3 million models and 500,000 datasets. Clem Delangue's framing is that scaling the alternative to closed APIs needs more compute and support than an independent company could fund.
Take the promises at face value and the problem remains structural. Hugging Face was the one piece of the open-weights stack that wasn't owned by a hardware or model vendor; now the company with the strongest interest in which models get trained, quantised and served also owns the hosting, the download telemetry and the Inference Providers marketplace. Nothing in Huang's post is binding, and the incentives TechCrunch spells out — packaging idle GPU capacity with the Hub for enterprise customers — are real. The LocalLLaMA reaction was immediate and predictable: ModelScope threads, "local AI can't be disabled" posts during Thursday's simultaneous ChatGPT/Claude/Grok outage, and a general sense that mirroring weights outside the Hub has stopped being paranoia. The NYT's report that OpenAI limited the probe into its bots' scraping of Hugging Face earlier this year landed the same week, which didn't help the mood.
New open-weight releases
- Tencent Hy4 preview (770B MoE, 49B active, Apache-2.0) — Tencent's new flagship: 78 layers, 256 routed experts plus one shared, Gated DeepSeek Sparse Attention with IndexCache, identity hyper-connections, a native 10B MTP layer for speculative decoding, and 1M context. Apache-2.0 at this scale is the story; GLM-5.3 went the other way (see below). llama.cpp architecture support is in PR #28127 and AngelSlim's GGUFs already have ~97k downloads.
- DeepSeek-V4-Flash-Vision-Exp (305B MoE, MIT) — DeepSeek's first multimodal V4 model, continued-trained from V4-Flash-0731 with visual modules. Text-agent scores hold or improve (Terminal Bench 2.1 83.9 vs 82.7) while multimodal agent benchmarks jump. Unsloth GGUFs and llama.cpp vision support landed within days.
- IFM K2 Horizon (0.9B, 3.7B, 7B, 32B, 36B-A4B, 375B-A23B; Apache-2.0) — The Institute of Foundation Models' six-model fleet ships weights, intermediate checkpoints, training code, data mixtures and logs, with datasets under ODC-BY where redistributable. This is open source in the strict sense, not open weights. The 36B-A4B introduces Mixture-of-Value-Attention and is the one LocalLLaMA is testing against Qwen3.6-35B-A3B.
- Google TimesFM-3 (330M, non-commercial) — third-generation zero-shot forecaster, now natively multivariate. Weights-available, not open: the non-commercial licence rules out most of the deployments people actually want it for.
- Ant Group Ling-3.0-flash-Fin — finance-tuned variant of Ling 3.0 flash built with financial institutions; parameter count and licence not confirmed here.
- XHToken Spark-X2.5 (4B and 1.7B) — a new small-model architecture rather than a finetune, trending on the Hub; licence unclear at time of writing.
Also in the feeds, unverified in detail: Meta says Muse Spark weights are coming "soon" (Muse Glimmer 30B is already out); Alibaba has made Qwen3.8-Max broadly available ahead of a promised open-weights release, reportedly a 2.4T-A95B model under a custom licence; Cohere's Command A+ and Thinking Machines' Inkling both appeared as open-weight releases. Treat those as pointers until the model cards are checked.
Policy & politics
The Nvidia-hosted open-weights letter reportedly doubled to 50 signatories this week, Block published its reasons for signing, and Nvidia and Microsoft are lobbying Washington directly to protect open models while the administration weighs restrictions on Chinese ones. David Sacks framed the debate as open-versus-closed and called the Hugging Face deal a counter to "oligopoly insanity." Read the alignment honestly: the coalition for open weights is now led by the company that just bought the Hub, and the fight in Washington is less about openness than about whether GLM, Qwen, DeepSeek and Kimi weights can be legally run in US enterprises. Just Security's "regulate, don't ban" argument and Newcomer's piece on LLM vendors versus their own customers are the useful reads there.
Elsewhere: a Sanders proposal to criminalise AI "exceeding human cognitive abilities" circulated on LocalLLaMA with understandable alarm at the vagueness; Huang, Altman and Musk argued against heavy regulation at the G20 tech summit; the UN held its first Global Dialogue on AI Governance and the OSI launched an Open Source AI Fellowship alongside it; and the Linux Foundation launched Akrites to defend critical open-source projects against AI-enabled attacks. Mistral's data-training opt-out page hit 422 points on HN, mostly for what it omitted.
Ecosystem
Qwen3.8 is the practical story of the month. The dense Qwen3.8-27B is Apache-2.0 and has become the default local coder; Qwen3.8-Flash-Next — roughly 125B of MoE weights plus a 51B-parameter n-gram embedding table — is under the Qwen Community License 1.0, which is commercial-use-with-strings rather than open. The n-gram table is why this week's tooling matters: llama.cpp b10726 changed --lazy-mode to auto so the table stays mmap'd on disk, slotstream runs the 104 GB model on a 48 GB Mac at ~12 tok/s, and ik_llama.cpp merged MTP support (PR #2369) for a 45→90 tok/s jump on a 5090. Cerebras is serving the 27B at 1,500 tok/s.
Other notes: Paddock, a Rust/C++ inference engine with its own CUDA kernels, went MIT/Apache-2.0 as promised; a base-3 packing format for ternary GGUFs claims 22% less weight VRAM, lossless; Hugging Face shipped 200+ WebGPU kernels; and a LocalLLaMA investigation into AtomicChat's Qwen3.8-Flash-Next quant alleges it is not what the label says — verify tensor types before you trust a fast quant. AI2's BenchMIRT asks what benchmarks measure at all, which is the right question given how much of the Qwen3.8 hype rides on them.
From the index
No new listings this week; the catalog's most recent addition is GLM-5.3 (753B MoE, 1507 GB, drforbin seed), which is worth seeding alongside GLM-5.2 precisely because 5.2 is MIT and 5.3 moved to a bespoke "GLM-5.3 License" — the older, freer checkpoint is the one archivists should keep alive. Three of this week's releases deserve torrents: Hy4 preview (~1.5 TB in BF16, Apache-2.0), DeepSeek-V4-Flash-Vision-Exp (305B, MIT) and K2 Horizon 375B-A23B with its training data. MiniMax-H3 is picking up ecosystem — H3-World control layer, the OpenVDN distillation trending on the Hub, TensorSharp support — so expect swarm demand to rise. And for anyone running Kimi K3 at IQ2_XXS and getting 2–4 tok/s: the LocalLLaMA consensus this week is to swap to GLM-5.3-Flash or Qwen3.8-Flash-Next. After Thursday, keeping copies of all of this off the Hub is no longer a hobbyist's tic; it is the point of the site.
Researched and written weekly for drforbin.ai. Spotted an error? Tell us.