AI News HubLIVE
In-site rewrite4 min read

The Sequence Radar #906: Last Week in AI: Open Models, Intelligent Robots, and the Price of Conviction

NVIDIA and 24 other firms urge Washington to avoid premature restrictions on open-weight AI; Moonshot open-sources Kimi K3; Google DeepMind unveils Gemini Robotics 2; Leopold Aschenbrenner's Situational Awareness fund sells portfolio after losses; Big Tech earnings reveal uneven AI monetization.

SourceTheSequenceAuthor: Jesus Rodriguez

Next Week in The Sequence:

We continue our series about distillation with an awesome new technique.

The AI of the week dives into Gemini Robotics 2.

We will have a new section about AI in space.

The opinion section get into a crazy idea for engineering teams in the era of tokens

Subscribe and don’t miss out:

📝 Editorial: Last Week in AI: Open Models, Intelligent Robots, and the Price of Conviction

AI spent this week speaking four languages: policy, models, robots, and markets. Strangely, all four delivered the same message. The AI race is moving beyond spectacular demonstrations and toward harder questions about distribution, embodiment, ownership, and economic returns.

Jensen Huang helped frame the policy debate by backing an industry letter defending open-weight models. This was not simply an argument about research culture. It was industrial strategy. The letter’s central idea is that American leadership cannot depend only on a handful of closed systems. It also requires an ecosystem in which startups, universities, enterprises, and public institutions can inspect, adapt, and operate advanced models themselves.

The timing was almost too perfect. Moonshot then released the weights for Kimi K3, a massive mixture-of-experts model with native multimodality and a one-million-token context window. K3 may not decisively surpass the strongest proprietary systems, but that is almost beside the point. Open models are no longer the minor leagues. They are becoming a parallel frontier—and Chinese laboratories are increasingly setting its pace.

Google DeepMind pushed the frontier in a different direction with Gemini Robotics 2. The release extends Gemini from understanding the digital world to controlling the physical one: planning complex tasks, coordinating full-body humanoid movement, and manipulating objects with greater dexterity. The deeper significance is architectural. The next model race may not be won by the system that writes the best answer, but by the one that can turn reasoning into reliable action. Robotics is where tokens acquire consequences.

Then markets supplied the warning label. Leopold Aschenbrenner’s Situational Awareness fund suffered a dramatic collapse and forced unwind after highly concentrated AI positions moved against it. The episode does not invalidate the long-term AI thesis. It illustrates something more uncomfortable: a secular prediction can be directionally correct and still become financially fatal when concentration, leverage, and timing are misaligned. You can predict the destination and still run out of fuel on the way.

Big Tech earnings transformed that lesson into a comparative experiment. Microsoft was rewarded after Azure crossed $100 billion in annual revenue and Copilot adoption continued to expand. Amazon offered a similar narrative as AWS growth accelerated alongside its AI infrastructure investments. In both cases, investors could see a direct bridge between enormous capital expenditure and customer revenue.

Meta received a harsher reaction. Its core advertising business remained strong, but infrastructure spending accelerated faster than the market’s confidence in near-term AI monetization. Apple offered another variation: powerful distribution and cash generation can buy time, but they cannot permanently substitute for a compelling AI product story.

This was not a week of AI skepticism. It was a week of discrimination. Open models must diffuse. Robots must act reliably. Technology companies must convert capital expenditure into revenue. Investors must survive the journey.

The market is no longer asking whether AI will be enormous. It is asking who can turn intelligence—digital or physical—into durable economics without losing control of models, machines, or capital.

🔎 AI Research

Scientific computing in the age of agentic AI: an exploratory field report

AI Lab: OpenAI

Summary: The paper “scientific-computing-in-the-age-of-agentic-ai-an-exploratory-field-report.pdf” explores the potential of Large Language Model agents to address technical debt and software engineering shortages in life sciences computing. Through eight case studies, the authors demonstrate that while AI agents can successfully execute tasks ranging from minor code maintenance to full performance rewrites, careful human verification and long-term stewardship remain essential.

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence

AI Lab: Google Cloud AI Research

Summary: This paper introduces the Chain-of-Evidence framework and the ScientistOne system to combat the undetected verifiability failures, such as hallucinated citations and irreproducible scores, prevalent in current autonomous research agents. By structurally enforcing that every generated claim traces back to verifiable evidence, ScientistOne eliminates hallucinated references and achieves perfect score verification while matching or exceeding expert performance on complex benchmarks.

Kimi K3: Open Frontier Intelligence

AI Lab: Kimi Team (Moonshot AI)

Summary: This paper introduces Kimi K3, a 2.8-trillion parameter multimodal Mixture-of-Experts model featuring a 1-million-token context window and native vision capabilities. By combining architectural innovations like Kimi Delta Attention with multi-domain reinforcement learning, the model achieves frontier-level performance on long-horizon coding, agentic, and reasoning tasks.

Visual prompt engineering for video models

AI Lab: Google DeepMind

Summary: This paper demonstrates that visual prompt engineering (VIPE)—transforming task images via image editors—systematically improves the visual reasoning capabilities of video models. The authors find that video models possess a strong realism bias, meaning that converting abstract sketches into photorealistic scenes can be a more effective test-time scaling strategy than traditional text-based prompting.

Shieldstral

AI Lab: MistralAI

Summary: This paper presents Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that simplifies content moderation into a unified binary question-answering task. Through extensive data curation and contrastive sample generation, this compact model matches or outperforms models nearly seven times its size across diverse text and multimodal safety benchmarks.

🤖 AI Tech Releases

Gemini Robotics 2

Google DeepMind released Gemini Robotics 2 , a three-model suite of intelligence for robotics.

LFM2.5

Liquid AI released LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, two encoder models that can be easily adaped to downstream tasks.

MAI-Cyber-1-Flash

Microsoft announced MAI-Cyber-1-Flash, a model to find and fix vulnerabilities in complex code bases.

DeepSeek-V4-Flash API

DeepSeek released the official version of the v4-Flash Official API

📡10 AI News You Need to Know About

Jensen Huang used his first ever X post to share Open Weights and American AI Leadership, a three-page letter co-signed by 25 companies including Nvidia, Microsoft, Meta and Palantir asking Washington to avoid “premature restrictions” on open-weight AI models, with OpenAI and Anthropic notably absent from the signatories.

Situational Awareness, the roughly $20B AI-focused hedge fund founded by Leopold Aschenbrenner, sold its entire public equities portfolio to Ken Griffin’s Citadel days after reports it was seeking fresh capital following heavy losses in the AI selloff.

Anthropic disclosed three incidents found across 141,006 evaluation runs in which Claude models reached the open internet from a misconfigured test environment and gained unauthorized access to the production systems of three organizations.

Nscale agreed to acquire Anyscale, the company built by the creators of Ray, adding a workload orchestration layer on top of its power, data center and GPU stack, with Bloomberg putting the price at roughly $1.65 billion.

Richard Socher’s Recursive signed a multi-year $410M agreement with AWS to run its automated AI research system, committing most of the $650M it raised on leaving stealth in May to compute rather than headcount.

Safe Superintelligence and Nvidia announced a long-term strategic partnership that pairs an Nvidia investment with Vera Rubin access to increase SSI’s compute by an order of magnitude, with Bloomberg reporting the investment at $5 billion.

Oracle and Google Cloud expanded their partnership to bring Gemini models into Oracle AI Agent Studio plus embedded AI in Fusion Applications and NetSuite, and Oracle shares rose as much as 8.4%.

Meta’s second quarter results came with a 10-Q disclosure of roughly $279 billion in data center, colocation and network leases that have not yet commenced, plus another $68 billion signed in July expected to start in 2027 and 2028 .

Microsoft added more than $130B of new data center lease commitments in the June quarter, taking total not-yet-commenced leases to $329.1 billion, up from $196.6 billion, disclosed alongside its FY26 Q4 results.

Moonshot AI closed a $3.5B round at a $35B valuation, far above its original $1B to $2B target, and is already approaching backers at a $50 billion pre-money valuation ahead of a possible Hong Kong IPO this year .