AI News HubLIVE
In-site rewrite5 min read

The Sequence Radar #897: Last Week in AI: China, Compression and the Open-Model Race

This week's developments in AI shift focus from raw scale to distribution and openness. Highlights include Thinking Machines' open-weight Inkling, Moonshot AI's 2.8T parameter Kimi K3, PrismML's phone-runnable Bonsai 27B, OpenAI's self-play red-teaming system GPT-Red, and Xi Jinping's call for open-source AI as a global public good at the World AI Conference in Shanghai.

SourceTheSequenceAuthor: Jesus Rodriguez

Next Week in The Sequence:

We continue our series about model distillation techniques.

In the AI of the Week , we discuss Thinking Machine first open weights model.

The opinion section reviews the idea that Google’s is by far the biggest threat to NVIDIA’s dominance.

Subscribe and don’t miss out:

📝 Editorial: Last Week in AI: China, Compression and the Open-Model Race

For years, AI progress has been narrated as a horse race: larger models, higher benchmark scores, more expensive clusters. This week broke that frame. The most consequential developments were not simply advances in intelligence; they were competing answers to a deeper question: Who gets to possess, adapt and govern it?

Thinking Machines Lab entered the model arena with Inkling, a 975-billion-parameter mixture-of-experts model with 41 billion active parameters, native text, image and audio inputs, a one-million-token context window, open weights and an Apache 2.0 license. The company’s emphasis on calibration, controllable reasoning effort and customization is more interesting than any leaderboard placement. Inkling is a wager that the next frontier will be models people can shape around their own judgment—not merely rent from a centralized oracle.

Moonshot AI pushed in the opposite direction on scale while converging on the same politics of access. Kimi K3 carries 2.8 trillion parameters, activates 16 of 896 experts, supports vision and a million-token context, and is aimed at long-horizon coding and knowledge work. Moonshot describes it as the first open model in the three-trillion-parameter class, although the full weights are not due until later this month. That caveat matters, but so does the trajectory: Chinese labs are no longer competing only on cost. They are using openness itself as a strategic lever.

PrismML’s Bonsai 27B makes the week’s most radical claim through compression rather than scale. Its ternary model is 5.9GB, while the one-bit variant is 3.9GB—small enough, PrismML says, to fit within a modern smartphone’s usable memory. The company reports retaining most of the performance of the full-precision baseline across its benchmark suite. Independent testing will determine how durable those numbers are. Yet the direction is unmistakable: intelligence density may become as strategically important as raw intelligence. Local models alter latency, privacy, cost and even the bargaining power between users and cloud providers.

A different type of model was also introduced this week by OpenAI. GPT-Red introduced a different—but equally consequential—form of scaling. It is not a consumer model or a new chatbot, but an internal automated red-teaming system trained through self-play to attack other models, observe their defenses and invent progressively stronger prompt injections. In unfamiliar test environments, GPT-Red successfully compromised GPT-5.1 in 84% of scenarios, compared with 13% for human red-teamers, and its attacks were subsequently used to make GPT-5.6 substantially more robust. The deeper idea is a new scaling law for safety: as models become more capable, the systems testing them can scale alongside them. The future may not be defined by models that simply improve themselves, but by machine-speed ecosystems in which attackers and defenders continuously co-evolve.

The common denominator is not openness as an ethical abstraction, but distribution as an engineering objective. Open weights permit adaptation; sparse architectures make colossal models economical; extreme quantization brings inference to the edge. Each move weakens a different bottleneck that has kept advanced AI concentrated in a few institutions.

That technological story now has an explicitly geopolitical counterpart. At Shanghai’s World AI Conference, Xi Jinping cast open-source AI as a global public good, promoted Chinese support for developing countries and elevated a new international AI cooperation organization as a vehicle for shaping global governance. The rhetoric of access is not geopolitically neutral; standards, training programs and model ecosystems can build dependencies as effectively as chips and cloud infrastructure.

Taken together, these announcements describe a frontier that is fragmenting—and perhaps democratizing. The decisive contest is no longer just over who can build the smartest system. It is over whether intelligence will be centralized or portable, proprietary or adaptable, governed by firms, states or communities. This week, AI stopped looking like a single race. It began to look like an emerging world order.

🔎 AI Research

GPT-Red: Unlocking Self-Improvement for Robustness

AI Lab: OpenAI

Summary: This article introduces GPT-Red, an automated red-teaming model that uses self-play reinforcement learning to efficiently discover vulnerabilities and prompt injection attacks. By integrating GPT-Red into their training pipeline, researchers successfully enhanced the robustness of production models like GPT-5.6 Sol against malicious instructions without degrading their core capabilities.

A Theory of Contrastive Learning with Natural Images

AI Lab: CSAIL, MIT and Hebrew University of Jerusalem

Summary: This paper analytically derives the optimal representations for contrastive learning with natural images, proving that simple augmentations optimally result in a partial whitening process computable by a basic CNN with sinusoidal filters. Experimental results confirm that CNNs trained with contrastive loss on various datasets naturally learn these sinusoidal filters and perform partial whitening as predicted by the theory.

On the Interpolation Effect of Score Smoothing in Diffusion Models

AI Lab: Google Research

Summary: This paper investigates the generative creativity of diffusion models, proposing that their neural network backbones naturally learn a smoothed version of the empirical score function rather than memorizing the exact training data. Through theoretical analysis and numerical experiments, the author demonstrates that this score smoothing guides the denoising dynamics to generate novel samples that smoothly interpolate between training points, even across complex nonlinear manifolds.

Video Generation Models are General-Purpose Vision Learners

AI Lab: Google DeepMind

Summary: The authors propose GenCeption, a unified generalist vision model that leverages a pre-trained text-to-video diffusion backbone to perform a wide variety of dense and sparse visual tasks via text instructions. By repurposing iterative diffusion into an efficient feed-forward architecture, GenCeption achieves state-of-the-art performance across diverse perception tasks while demonstrating emergent behaviors like sim-to-real transfer and zero-shot generalization.

Metacognition in LLMs: Foundations, Progress, and Opportunities

AI Lab: Yale University and University of California, Irvine

Summary: This paper presents a comprehensive taxonomy and review of metacognition in Large Language Models, detailing how models’ abilities to monitor and regulate their own cognitive processes are measured, elicited, and improved. It synthesizes current findings on metacognitive capabilities like confidence calibration, introspection, and knowledge boundary detection, highlighting the potential of these mechanisms to enhance LLM reliability, reasoning, and human-AI collaboration.

ADVANCED MATHBENCH: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

AI Lab: Shanghai AI Laboratory

Summary: ADVANCED MATHBENCH introduces a rigorous evaluation suite focusing on the generation and process-level verification of advanced, natural-language mathematical proofs at the undergraduate and doctoral qualifying-exam levels. The benchmark reveals that frontier LLMs still struggle significantly with constructing and verifying complex mathematical proofs, highlighting a critical bottleneck in their ability to accurately detect subtle logical errors.

🤖 AI Tech Releases

Inkling

Thinking Machines open sourced Inkling, a large MoE model with a 1M context window.

Kimi K3

Moonshot AI released Kimi K3, a massive 2.8T parameter model optimized for long horizon coding, reasoning and kowledge work.

Robostral Navigate

Mistral released its first model optimized for embodied AI.

Bonsai 27B

PrismML released Bonsai 27B, a model that can run entirely on a phone.

📡10 AI News You Need to Know About

Xi Jinping made his debut at the World AI Conference in Shanghai touting China’s low-cost AI and calling for an open, cooperative technological order, saying AI development should be a symphony of international cooperation rather than a solo performance by one country.

Reuters reported that China’s Cyberspace Administration approved Apple Intelligence for launch on the back of a deal integrating Alibaba’s Qwen models into iOS, iPadOS, macOS, and visionOS, with Baidu also confirming it is working with Apple on features for Chinese users.

Emergent announced a $130 million Series C led by Creaegis at a $1.5 billion valuation, a fivefold jump in months, on the back of a $120 million revenue run rate and 200,000-plus paying customers for its AI software creation platform.

Demis Hassabis published a proposal on X calling for a FINRA-style independent standards body, funded by industry and backed by the US government, that would review frontier models up to 30 days before release and eventually gate deployment in the US market. (Replaced TechCrunch with Hassabis’s original post.)

Reflection AI signed a $1 billion-plus compute deal with Nebius running through 2029 that gives the open-model lab access to Nvidia’s GB300 chips, its second major capacity grab after last month’s SpaceX agreement.

Bloomberg reported that Google is months behind schedule on Gemini 3.5 Pro because the model is falling short of internal goals, especially in coding, frustrating employees who worry Anthropic and OpenAI are pulling ahead.

Greylock announced Greylock 18 on X, a $1.5 billion early-stage fund, its eighteenth, aimed at concentrated bets on AI-native founders, with managing partner Asheem Chandna arguing the next trillion-dollar companies are still ahead.

TSMC’s June revenue report showed sales of NT$442.68 billion for the month, up 67.9% year over year, lifting June-quarter revenue 36% to roughly $39.6 billion and confirming AI demand remains intact. (Replaced Bloomberg with TSMC’s own monthly revenue release.)

Walden Robotics launched out of stealth, a Toyota Research Institute spinout led by MIT’s Russ Tedrake, with roughly $300 million in seed funding at a $1.1 billion valuation co-led by Toyota and Deviation Capital, and its general-purpose robots already working production shifts at a North American Toyota plant.

SK Hynix raised $26.5 billion by selling 177.9 million ADRs at $149 each in the largest-ever US listing by a foreign company, topping Alibaba’s 2014 record, just as Commerce Secretary Lutnick pressed the memory maker to build new US fabs.