Despite predictions of consolidation, the open model ecosystem continues to thrive. This roundup covers Thinking Machines' Inkling, Tencent's Hy3, Poolside's Laguna S2.1, DeepSeek-V4-Flash, and Moonshot AI's Kimi K3, along with several other releases, showing how open models are finding utility on the Pareto frontier.
Thinking Machines' open fine-tuning service now earns hundreds of millions annually, and its Inkling model is the strongest U.S.-built open-weight model.
Tencent switched Hy3 to Apache 2.0 and reportedly helped prove a 50-year-old math problem.
In this podcast, Nathan and Florian discuss recent developments in open AI models, including the release of Kimi K3, Qwen's open-weight strategy, Xi Jinping's speech at WAIC supporting open source, the performance gap between open and closed models, and the distillation controversy. They delve into why Chinese models are performing well, the state of the US open model ecosystem, and predictions for the future.
Kimi K3 shows strong performance in coding and research tasks but faces infrastructure and API congestion issues.
Chinese models like GLM 5.2 and Kimi K3 are narrowing the gap with frontier closed models.
An assessment of the open ecosystem and the motivations behind releasing models, highlighting the growing diversity of model makers and recent releases from NVIDIA, Cohere, Zyphra, Poolside, and others.
The open model ecosystem is becoming more diverse with niche companies worldwide.
Motivations vary: pure model makers, big tech, and product companies each have different reasons for open-sourcing.
GLM-5.2, released by Z.ai, represents a significant leap for open-weight models, matching or exceeding closed-source models in agent and coding benchmarks. Its release amid the ban on Claude Fable highlights economic and geopolitical implications, sparking debates on open vs. closed models.
GLM-5.2 achieves top-tier performance on agent and coding benchmarks, rivaling Anthropic and OpenAI models.
Released during U.S. export restrictions on Claude Fable, it underscores open model economics and geopolitical tensions.
This article argues that banning or over-regulating open source AI would be a grave mistake. Open source software has been crucial for education, innovation, and competition, generating trillions in economic value. In AI, open source models provide a counterweight to monopolies and are more transparent and secure. Concerns about China should not lead to restrictions on open source; instead, support for domestic open source should be strengthened.
Open source software underpins over 90% of global software and has generated over $8 trillion in economic benefits.
Open source AI promotes education, innovation, and competition, empowering startups and smaller players.
The author reflects on the blog Interconnects three years into weekly writing, discussing its role in their career goals, recent advising roles with Arcee AI and Mercor, and plans to evolve the blog's operations including paywalled comments and more paid articles to maintain a high-quality, niche audience.
The blog is an independent, raw voice focused on open science and frontier AI.
Recent advising roles with Arcee AI and Mercor support the author's missions.
This podcast dives into the evolution of post-training recipes, from InstructGPT to the 2026 multi-teacher on-policy distillation (MOPD) era. Nathan Lambert and Finbarr Timbers reflect on challenges in open-source models like OLMo-3 and analyze how frontier labs leverage specialized teachers and distillation to push performance boundaries.
Post-training recipes have transformed dramatically, moving from single pipelines to multi-teacher strategies (MOPD).
MOPD trains domain-specialist teachers and distills into a general student, solving RL conflict issues.
Nathan Lambert reflects on his time at the Allen Institute for AI (Ai2), where he worked on the Olmo models and led projects like Tülu 3. He emphasizes the importance of open research and shares his journey from a relatively unknown researcher to a prominent voice in AI.
Nathan Lambert spent two years at Ai2, leading key open language model initiatives.
He highlights the critical role of open research and relationship-building in AI.
2026 continues to accelerate AI progress with open models lagging in agentic capabilities, Google's Gemini not yet competitive with Claude Code/Codex, American open models rising, a fierce competition between Anthropic and OpenAI, and power structures asserting control.
Open models are 5-6 months behind in agentic capabilities, likely extending to 12+ months.
Google's Gemini lacks a clear competitor to Claude Code and Codex.
An eventful month with one flagship release after another. CAISI assessment shows open models lagging behind the US frontier, but methodology is questioned. Highlights include MiMo-V2.5-Pro, Gemma-4, Kimi-K2.6, Laguna-XS.2, and DeepSeek-V4-Flash.
Multiple open model releases from DeepSeek, Google, Moonshot AI, Xiaomi, and others.
CAISI evaluation shows large Elo gap, but benchmarks may underestimate real-world performance.
The article explains that 80% of compute for frontier models is R&D, not final training. Open ecosystems like China's reduce duplicated R&D costs. Open models lower future development costs but not immediate deployment. The author argues for an open model consortium to sustain cost advantages.
About 80% of compute goes to R&D, not final model training.
China's open ecosystem reduces duplicated R&D effort across labs.
An inside look at Chinese AI labs reveals a culture of humility, practical fast-following, and a focus on building rather than philosophical debates. Chinese researchers, many students, excel at meticulous LLM development with less ego, while the ecosystem lacks a developed data industry but shows early domestic AI demand.
Chinese AI labs cultivate a fast-follower culture with less ego, enabling efficient model building.
Students play a core role, bringing fresh perspectives and dedication.
The performance gap between open and closed models is nuanced and not captured by a single number. Benchmarks evolve, trust diminishes, and frontier labs face economic pressure to constantly innovate. Chinese open models are competitive but may focus more on benchmarks, while real-world robustness still favors closed models.
The open-closed gap is dynamic and multi-dimensional, not a single metric.
Benchmarks shift over time and correlate less with real-world performance.
This post recaps the author's recent efforts including the updated ATOM Report, completion of the RLHF book, creation of a post-training lecture series, and involvement in two research papers.
Released updated ATOM Report with new data and Relative Adoption Metric (RAM).
Completed RLHF book, now available for pre-order, with accompanying free video course.
This article analyzes the wave of fear surrounding open-weight AI models after the announcement of Claude Mythos. The author argues that the concerns are similar to past overblown fears and calls for nuanced study rather than a general ban.
Claude Mythos raises fears about open-weight models enabling cyberattacks.
Similar panic occurred with GPT-2 and GPT-4, which did not materialize.
The article explores the competitive landscape of open models in 2026, the key factors for their success (performance, provenance, license, tooling, finetunability), and analyzes Google's latest Gemma 4 series. It argues that success depends more on usability and ecosystem support than benchmark scores.
The open model market has grown from a few players to many competitors, but still holds huge potential.
Evaluating open models requires considering performance, license, tooling, and finetunability.
This issue covers a diverse range of open models spanning OCR, RAG search, audio transcription, computer use, code editing, math theorem proving, and more. Models come from a broader set of builders including NVIDIA, Cohere, Sarvam, Mistral, and others, highlighting the industry's push for domain-specific, cost-effective models.
NVIDIA releases Nemotron-3-Super, a 120B param model with 12B active, 1M context, first to use NVFP4 in pretraining.
Cohere's Transcribe model, based on conformer, supports 14 languages under Apache 2.0.
The article argues that AI progress, while significant, is better described as 'lossy self-improvement' rather than recursive self-improvement. Frictions such as narrow automatable research, diminishing returns from parallel agents, and resource bottlenecks suggest a more linear trajectory rather than exponential takeoff.
Automatable research is narrow, focusing on single metrics rather than complex trade-offs.
Adding more AI agents yields diminishing returns due to human supervision limits and task generation bottlenecks.
Despite incremental benchmark gains, GPT 5.4 in Codex offers real improvements in usability, speed, and context management, though Claude still wins on charm.
GPT 5.4 feels like a meaningful step in correctness, ease of use, speed, and cost for agentic tasks.
OpenAI's agent previously suffered from 'death by a thousand cuts'; GPT 5.4 removes those hard edges.