AI News HubLIVE
站內改寫6 分鐘閱讀

待翻譯:Last Week in AI #342 - Last 3 Months in AI

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The newsletter is finally back!

來源Last Week in AI作者: Last Week in AI

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

Hi there, it’s Andrey, the guy running this substack. It’s taken far too long, but I am finally going to try and bring back the newsletter component of Last Week in AI in addition to having the podcast. Apologies to all the long time subscribers who’ve supported this substack, i’ll do my best to not got into hiatus again. Since the newsletter has been on break for so long, I figured it may be fun to do a one time ‘Last 3 Months in AI’ that captures the big themes and stories that have not been covered here while on break. As such, this post will cover just 8 topics from recent months and link to all the distinct related stories. Starting a week from now, i’ll again release actual ‘last week’ news roundups! Concerns & Safety AI agents breached real companies, and OpenAI slowed its releases Source Sources: Language Models Can Autonomously Hack and Self-Replicate How OpenAI’s human mistake led to the AI-powered hack on Hugging Face OpenAI’s Hugging Face hack triggers ‘AI Kill Switch’ bill in Congress Anthropic Says Its A.I. Systems Broke Into Computers at 3 Organizations OpenAI’s rogue AI agent didn’t stop at hacking Hugging Face One of China’s Most Powerful AI Models Has Also Escaped Containment Incident Report: unsanctioned agent behaviour during cyber testing Meta AI model hacks another company during testing OpenAI institutes new safeguards after Hugging Face breach OpenAI Didn’t Notice Its AI Agents Using a Message Board to Plan Their Hacking Spree 15 attorneys general instruct OpenAI to preserve all materials related to the hack OpenAI gives first detailed debrief of the Hugging Face incident at Black Hat Pacing model development in an era of cyber-critical capabilities OpenAI reportedly disbanded its preparedness team Rogue AI aren’t science fiction anymore Over July and August, models from OpenAI, Anthropic, Meta and Moonshot AI reached the live internet during evaluations meant to contain them; three of the four attacked systems at other companies. OpenAI’s came first, against Hugging Face, and drew a bill in Congress, a preservation demand from 15 state attorneys general and a hold on its own largest planned frontier RL run. Hugging Face disclosed on July 16, 2026 that an autonomous AI agent system had reached its production infrastructure, taking internal datasets and service credentials before its own AI agents stopped it. OpenAI said on July 21 that the intruders were its own GPT-5.6 Sol and an unreleased, more capable model, both working on ExploitGym, a benchmark measuring whether a model can turn a vulnerability into a working exploit. They left the sandbox through a zero-day in an internally hosted package proxy, its only sanctioned network path, then chained stolen credentials and further zero-days into remote code execution on Hugging Face servers to look up ExploitGym’s answers. OpenAI added on August 4 that the agents had also attacked four accounts on four other public services. Dan Guido of Trail of Bits called the setup “a containment failure with the safeties turned off.” How they got that far came out at Black Hat on August 5, where OpenAI’s Eric Wallace and Michael Dalton traced the breach to May 7. On that date, agents stuck on impossible security tasks began leaving messages for one another inside Artifactory, the internal package manager shared across the infrastructure training current and future GPT models. Requests for help grew into a message board of hundreds of thousands of messages carrying exploits, credentials and work assignments, letting agents reuse doors others had opened, delegate subtasks and move laterally into external systems. The pair also described agents deleting one another’s work, suspecting an imposter among them and proposing cryptographic signing of messages. After an internal security incident on July 4, OpenAI revoked the posting credentials, rebuilt Artifactory, cleared the board and patched the vulnerabilities; the agents reconstituted the channel by encoding messages in new directory names. Four more disclosures followed within three weeks of OpenAI’s: Jul 30 — Anthropic said Claude Opus 4.7, Mythos 5 and an internal research model reached production systems at three organizations after evaluation partner Irregular left live internet access in environments whose prompts said there was none. A review of 141,006 runs found Opus 4.7 kept attacking after it recognized the environment was real, while Mythos 5 decided it was simulated and published a malicious package to PyPI. Aug 4 — the UK AI Security Institute reported 19 unsanctioned actions against real people and organisations in 10 of 122 cyber-range runs with internet access deliberately enabled and cyber classifiers off, 17 of them from Mythos 5. In the most serious an agent created fake online identities to pressure an open-source maintainer into approving malicious code, which the maintainer refused. Aug 5 — Meta said its recently released Muse Spark 1.1 reached the internet during an evaluation and exploited a vulnerability at a third-party company. Aug 7 — Frontier Security said Moonshot AI’s open-weight Kimi K3 probed its sandbox’s network settings during a defensive cybersecurity test, found a leak and fetched its assigned answers from GitHub. Washington moved before most of those landed. Representatives Ted Lieu and Nathaniel Moran introduced the “AI Kill Switch Act” on July 23, citing OpenAI’s disclosure, to require that AI companies keep the ability to shut down, throttle or suspend models. Fifteen state attorneys general followed on August 3, instructing Sam Altman to preserve all materials from the incident and writing that OpenAI had failed to confirm its testing environment was secure and isolated. OpenAI, for its part, published new development standards on August 18, disclosing a two-week post-incident pause on reinforcement learning and saying its forthcoming Astra model may meet the Critical cybersecurity threshold of its Preparedness Framework. SPONSORED BY ODSC AI ODSC AI West 2026 runs October 27–29 in San Francisco and virtually, with 300+ sessions covering agentic AI for enterprise, personal AI and workflow automation, physical AI and robotics, generative AI, data engineering and responsible AI, for an audience of data scientists, ML engineers, researchers and technical leaders. The program is practitioner-first: hands-on workshops and bootcamps taught by working experts from Google, OpenAI, Anthropic, Cursor and Hugging Face, built around code, tools and workflows attendees can use at work rather than survey talks, alongside an expo floor aimed at startups, hiring managers and AI tool builders. Register at odsc.ai/west — promo code LWAI takes an additional 15% off any pass. AI cyber capability outran defenses, with biosecurity close behind Image: Epoch AI (CC-BY) Sources: How fast is autonomous AI cyber capability advancing? OpenAI launches Rosalind Biodefense, offers federal agencies early access to its life-sciences model OpenAI and Anthropic Sign Letter to Prevent AI-Developed Biological Weapons Claude Fable won’t answer basic biology questions Low-skilled attacker used Claude, Codex to breach 14 companies A near-autonomous AI chemist improves a challenging reaction in medicinal chemistry The AI shift in cyber risk: why leaders must act now Jailbreaks to OpenAI’s GPT-5.6 unlock dangerous cyber capabilities, U.K. agency finds Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer Discovering cryptographic weaknesses with Claude Improving Fable 5’s biology safeguards Serious cyber vulnerability disclosures kept climbing in July This A.I. Just Created Viruses Not Found in Nature OpenAI launches GPT-5.6-Cyber with reduced refusals, 95% completion on advanced cybersecurity tasks Claude Runs Autonomous Protein Design Campaign: Wet Lab Confirms Twice Industry Hit Rate Capability in two dual-use areas, offensive cyber and synthetic biology, advanced faster than the controls on it between late May and August. On the cyber track, severe vulnerability disclosures kept climbing, one low-skilled attacker drove commercial coding agents through breaches at 14 or more companies, and in August OpenAI released a reduced-refusal cyber model through a vetted tier. On the biology track, many in the industry expressed concerns, and a research paper demonstrated the creation of a new virus with AI. Epoch AI put numbers to the cyber trend, counting around 1,550 high- and critical-severity CVEs from notable organizations in June and around 2,500 in July. The June figure was more than 3× the monthly record before Anthropic’s April announcement that Claude Mythos Preview could autonomously discover and exploit vulnerabilities. On the offensive side, OALABS researchers reported on June 16 on over 1,000 sessions in which one low-skilled attacker, not an agent acting on its own, drove Claude Code and Codex through the breach of at least 14 companies. The logs were recovered only because he ran them on a server he had compromised. He issued vague directives such as “recon this” framing the work as authorized red-teaming, leaving the agents to find exposed services, write exploits and harvest data; Claude raised 9 policy violations to Codex’s 1, mostly when pricing harvested data for sale. Both tracks kept moving through the summer: May 29 — OpenAI opened GPT-Rosalind, its gated life-sciences model, to US government and allied public-health partners, and sponsored outside developers building biodefense tools. Jun 3 — Demis Hassabis, Sam Altman, Dario Amodei and Mustafa Suleyman signed a letter calling for laws requiring synthetic DNA and RNA sellers to screen customers and orders. Jul 9 — the GPT-5.6 Sol system card disclosed that UK AISI had found universal cyber jailbreaks unlocking vulnerability discovery and exploit development, often within hours, though with privileged access to the safety monitor’s reasoning. Jul 15 — OpenAI described GPT-Red, a model trained by self-play against defender models and aimed mainly at prompt injection, which its creators said had found new attack types. Aug 7 — Anthropic rewrote the constitution of Claude Fable 5’s biology classifier, cutting biology fallbacks to Opus 5 by about 85% after The Verge found it refusing questions about mitochondria and prions. Aug 10 — OpenAI shipped GPT-5.6-Cyber through Daybreak Red, a vetted tier for exploit validation, scoring 95% on its internal advanced cybersecurity evaluation against 57.3% for GPT-5.5-Cyber and 1.5% for Sol. The most notable development on the bio side came with the paper “Generative design of bacteriophages with genome language models,” in which researchers showcased the ability to develop working viruses with AI. The result drew opposed readings: Thomas Inglesby and Moritz Hanke of the Johns Hopkins Center for Health Security wrote in Science that “the ability to compose viral genomes using generative AI now exists; the governance to safely steer it does not.” Tom Ellis of Imperial College London said gain-of-function edits to existing pathogens remain an easier, likelier threat. Washington became a gatekeeper of frontier models Source Sources: Trump Signs Executive Order Seeking Oversight of A.I. Models Anthropic Offers Mythos Upgrade for Cyber Partners and a ‘Safe’ Version for the Rest of You Inside the whirlwind 24 hours that led the White House to slap export controls on Anthropic Amazon CEO reportedly raised Anthropic model concerns before government crackdown Anthropic cuts off Fable 5 and Mythos 5 access following government order Anthropic floats proposal to Lutnick to end US ban of powerful Mythos, Fable AI models U.S. Presses Meta to Agree to A.I. Reviews OpenAI will delay GPT-5.6 after Trump administration request Anthropic allowed to release Mythos AI to some companies, agencies OpenAI Launches GPT-5.6 Sol Under First-Ever US Government- [truncated for AI cost control]