跳到主要内容
AI News HubLIVE
站内改写6 分钟阅读

待翻译:The Sequence Radar - Issue 940: Last Week in AI: Opus 5.5 Gets Leaner, Meta Goes Wearable, Washington Talks to Beijing, and Claude Explores DNA

文章摘要

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Smarter models, wearable agents, scientific discoveries, and superpower diplomacy reveal how quickly AI is moving into the world.

来源TheSequence作者: Jesus Rodriguez
待翻译:The Sequence Radar - Issue 940: Last Week in AI: Opus 5.5 Gets Leaner, Meta Goes Wearable, Washington Talks to Beijing, and Claude Explores DNA
报告错误

纠错通道尚未开通,可先复制下方文章信息留存。

查看更正说明
直接读正文

AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。

Next Week in The Sequence: Another installment of our series about recursive self-improvement. We dive into Opus 5.5, DeepSeek’s amazing new paper about environments and Anthropic’s DNA discoveries. We discuss another interesting platform in robotics. We dive into the fascinating subject of AI compute as a new commodity, or is it? Subscribe and don’t miss out: 📝 Editorial: Last Week in AI: Opus 5.5 Gets Leaner, Meta Goes Wearable, Washington Talks to Beijing, and Claude Explores DNA A cheaper coding agent, glasses that can act on what you see, an unusual pattern in viral DNA, and two rival governments discussing AI risk. This week’s headlines read like four different newsletters. Together, they describe a shared transition: AI is moving deeper into the systems through which we work, discover, and make decisions. Anthropic’s Claude Opus 5.5 offers the most immediately practical example. The company reports stronger performance in coding and knowledge work, with typical workloads costing 40% less than Opus 5 and output generation running more than 30% faster. Those are Anthropic’s measurements, but the combination deserves attention. When an agent spends hours navigating a codebase, every unnecessary step costs money and introduces another opportunity for failure. Better economics expand the set of tasks worth delegating. A migration that once looked prohibitively expensive can become a routine overnight job—provided the results survive review. At Meta Connect, the emphasis shifted toward where those agents will live. Meta announced plans to bring Muse to its AI glasses, expanded its connections to shopping and productivity services, and previewed the pocket-sized Muse Charm. New audio glasses and lightweight VR glasses rounded out the hardware push. My read: Meta is betting that distribution and context will become decisive advantages. An agent that can see the object you mean requires less explanation. One connected to the services you use can turn that understanding into action. The product challenge is making those interactions reliable enough to become habits. That widening reach also explains why US–China AI talks matter. Washington proposed a mechanism for notifying Beijing about significant AI incidents, while both leaders identified AI as an area for cooperation during this week’s summit. The public record still offers limited evidence of concrete safeguards. Yet even a modest communication channel would address a real problem: rivals need ways to distinguish an accident from an attack. Competition can accelerate capability while simultaneously increasing the value of coordination. Building that coordination will require more than agreement that AI is consequential. The week’s most intriguing scientific announcement came from Anthropic’s biology lab. The company reported that Claude agents identified a previously uncharacterized enzyme system with CRISPR-like repeat structures. Roughly 950 agents searched for 21 hours, with human scientists performing the laboratory experiments. The underlying enzyme had appeared in earlier research; Claude identified previously unnoticed features around it. Its biological function and practical utility remain under investigation. That distinction matters. We have an early discovery, with substantial work ahead before anyone can call it a useful gene-editing technology. Still, the workflow is compelling: search broadly, identify anomalies, challenge hypotheses, and bring the strongest candidates to the laboratory. For me, the thread connecting these developments is the growing importance of what surrounds a model. Costs determine which jobs we delegate. Interfaces determine when we ask for help. Experiments determine whether a proposed discovery survives. Diplomacy determines how countries respond when something goes wrong. As intelligence becomes more accessible, those surrounding systems will increasingly determine how much value it creates. 🔎 AI Research DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale AI Lab: DeepSeek-AI, Tsinghua University Summary: DSec is DeepSeek’s production sandbox platform—unified FnCall/container/microVM/full-VM backends, EROFS composable layers, and 3FS on-demand image loading, co-designed with the RL loop for stateful rollouts under GPU preemption. A ~160-node unit serves ~3M sandboxes/day at >380K concurrency and >5K creates/sec; on-demand loading finishes an 8,192-container burst in ~35 min vs >60 min for eager pulls (1.71×) while cutting disk writes ~57%, and virtio-pmem+DAX trims peak host memory 40.2%. Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents AI Lab: Salesforce AI Research Summary: Instead of distilling trajectories into fixed write-time artifacts, JIT Mem stores raw episodes and trains a GRPO curator to synthesize a compact, task-conditioned payload at read time from immediate task success. It beats the strongest write-time baselines by +16.2 / +16.3 / +3.9 SR points on ALFWorld, WebShop, and τ²-bench, with an untrained Gemini curator already at 61.0 vs SkillOS 41.0 on WebShop, while cutting input tokens 50.3–56.3% and executor steps 28.4–31.4%. Schrödinger’s Code Repository: Have LLMs Learned SWE-bench or Memorized It? AI Lab: Shanghai Jiao Tong University, Xi’an Jiaotong University, East China Normal University Summary: SchrodingerRepo treats the SWE-bench repository as an evaluation-time latent variable—problem rewrite, namespace remap, layout reorder, and functionality-preserving rewrite—preserving executability while eroding memorized cues. Full transforms drop Pass@1 by 6.0–14.4 pp across GPT-5.1, GPT-5.4-mini, DeepSeek-v4-Flash, and Gemini-3.1-Flash-Lite (e.g. 46.8%→35.6% for GPT-5.4-mini), with 81.6–83.6% of extra actions spent on exploration and input tokens rising >2.5×; human checks find >65% of Verified instances show leakage evidence. The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks AI Lab: Microsoft, City University of Hong Kong Summary: Taste-Bench mines 502 decision-fork questions from engineering and research agent trajectories so models must pick the better branch before later outcomes are revealed. The best frontier model (GPT-5.6 Sol) scores only 59.7%, longer horizons are harder, and more reasoning budget does not help; distilling a privileged teacher into Qwen3.6-27B raises held-out taste accuracy 30.0%→47.9% (+17.9 pp) and lifts a fixed executor on 41 held-out SWE-bench Pro tasks from 14.6% to 33.7% success with student advice. Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms AI Lab: Alibaba Token Foundry, Georgia Tech Summary: VHD-Play reverses environment-first generation: sample and solve a mathematical mechanism first, then wrap its dynamics as stateful tools with the same reference for scoring, yielding 3,300 admitted environments at a few cents each. GRPO training lifts Qwen3.6-35B-A3B’s five-family mean agentic score from 0.204 to 0.815, transfers to eight unseen families, and on 365-day E-Commerce Bench reaches 3.4× the base ending balance while beating Qwen3.7-Max and gaining +2.84 on interaction-focused BFCL V4 cells. WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents AI Lab: Carnegie Mellon University, Tsinghua University Summary: WhatWorkedBench asks agents to submit a full predicted response surface after a measurement budget, scored against exhaustive CPU references for every conditional component effect across 36 tasks / 1,248 configs (35/36 tasks show sign-reversing interactions). Agents often pick good optima while missing effects; fitting a shared Gaussian process on the same observations raises effect recovery from 0.632 to 0.698 (Flash cohort) and from 0.303 to 0.455 on six added-family submissions, while encoding code equivalences lifts six-factor GP recovery from 0.248 to 0.462 at B=20. 🤖 AI Tech Releases Claude Opus 5.5 Anthropic released Claude Opus 5.5, the first model in the Claude 5.5 family—Fable 5.1-level on most work at about 40% lower cost than Opus 5 ($4/$20 per MTok input/output), live as claude-opus-5-5 on the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry, with preserved thinking and Mythos-class biology/cybersecurity safeguards. Muse on Meta AI glasses (Connect) At Connect 2026, Meta said Muse is coming to its AI glasses in the coming months—name-activate and act on what you see, plus expanded shopping/work connectors, Muse email, a pocket Muse Charm later this year, and Ray-Ban Meta Gen 3 / Ray-Ban Meta Audio as the lineup heads past 100 AI-glasses styles by year-end. 📡10 AI News You Need to Know About Meta is bringing Muse to its AI glasses with hands-free, look-and-ask control, plus real-time voice, expressive avatars, Mac computer use, a dedicated agent email, and more shopping connectors—rolling out on glasses in the coming months after Connect 2026. Cognition said it crossed $1 billion in annualized revenue run rate for Devin—less than two years after the coding agent became generally available—a milestone Bloomberg also covered. Enveda raised $311 million in a Series E led by Catalio Capital Management (new capital from Durable, ICONIQ, Lightspeed, Surveyor/Citadel, T. Rowe Price accounts, and others), bringing total funding above $845 million to push ENV-294, ENV-308, and ENV-6946 deeper into the clinic and scale its PRISM nature-chemistry discovery platform. Anthropic says Claude discovered ART—array-associated reverse transcriptases, a novel phage enzyme system with CRISPR-like DNA-repeat arrays—via ~950 agents scanning ~1.9B protein clusters over ~21 hours, with human scientists running the wet-lab follow-up in its new Bay Area biology lab (function still unknown; preprint out). Ema raised $77 million in a Series B led by Creaegis, with Accel, S32, and Prosus increasing stakes, bringing total funding to $140 million and more than quadrupling its valuation as it scales “AI Employees” that automate HR, IT, and finance workflows across enterprise apps. Snorkel AI raised $350 million at a $3.5 billion valuation in a Series E co-led by Insight Partners and S32 (Addition participating heavily), nearly tripling its prior mark to expand its agentic data factory for frontier-lab training data and environments. PrismML demoed its 1-bit Bonsai 2-billion-parameter vision-language model running locally on Qualcomm Snapdragon AR1 Gen 1 smart glasses at Snapdragon Summit, fitting about 4× more parameters into the same memory envelope as prior on-glasses models (no retail glasses using it announced yet). TechCrunch reported that vibe-coding startup Lovable has crossed about $600 million in annualized revenue (up from ~$500 million in June), with co-founder Fabian Hedin saying at HumanX that people at roughly two-thirds of Fortune 500 companies now use the platform—after an August Series C that valued the company at $13.3 billion. At the White House summit, Xi told Trump that China and the U.S.—both major AI powers—have more reason to cooperate than compete, urging continued AI dialogue on risks and benefits, joint work against misuse, and keeping AI under human control; Trump said AI “bears on the future of humanity” and that the two sides should keep talking, after trade teams held their first U.S.–China AI dialogue and floated an incident-notification channel. Oracle sent a force majeure notice on Project Jupiter, its New Mexico Stargate campus with Blue Owl (Bloomberg first reported), seeking to delay payments if the planned 2.45 GW site misses its 2028 online date rather than exit as tenant—while saying the project “remains on our planned schedule,” amid gas-pipeline and air-permit delays.

展开要点与分析

文章情报

工程师进阶

要点

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Smarter models, wearable agents, scientific discoveries, and superpower diplomacy reveal how quickly AI is moving into the world.

要点与分析由自动化流程生成,可能有误,请结合原始来源核实。