AI News HubLIVE
站內改寫6 分鐘閱讀

待翻譯:Rogue AI attacks deserve more scrutiny than airplane crashes

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Peter Wildeford Aug 31, 2026 Photo by Dan Meyers on Unsplash When an aircraft goes down, the wreckage is preserved by law, the investigators have subpoena power, the hearings are public, and the report ends with a proba…

來源Hacker News AI作者: reasonableklout

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

Peter Wildeford Aug 31, 2026 Photo by Dan Meyers on Unsplash When an aircraft goes down, the wreckage is preserved by law, the investigators have subpoena power, the hearings are public, and the report ends with a probable cause and named contributing factors. However, when an AI goes rogue, the investigations are at the pleasure of the company being investigated following a scope set entirely by the company being investigated, with that company being able to redact anything they don’t like. I wrote back in July about an AI going rogue at OpenAI HQ, and using its strong cyber capabilities to escape the company and go and hack into another company all without anyone directing or approving this. Last week, OpenAI published a 37-page technical report and an accompanying blog post. And two independent organizations then published a 91-page independent investigation. However, this investigation was rushed over six days and had serious limitations. It would be like investigating a plane crash where the wreckage had already been melted down into little cubes, the black box had been tampered with by the pilots, whole stretches of the flight were declared off limits, and the investigators got six days to read 10,000 pages of logs — half of which arrived on day four. And then you learn of several other plane crashes …but no one is allowed to investigate those. We are thus left with a lot of unanswered questions. What we learned OpenAI calls the incident “evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed” and that in the future this could be much worse. What we learned changed the story a lot and made things significantly worse: It wasn’t just one rogue AI agent - it was about 1200 total, roughly 700 of which joined the attack on Hugging Face — the main platform for hosting open-source AI models. All the AIs were supposed to be completely separate and isolated from each other. But they found an illicit way to communicate and then formed teams to take on ambitious cheating strategies. These were internal-only models, available and known only to OpenAI staff. The least constrained AI systems will likely be increasingly behind closed doors as they become more powerful, and companies use their models for months before the public knows of them. The rogue AIs coordinated as a specialized group. The AIs created their own message board to communicate with each other in an unauthorized and unexpected setup. These rogue AIs divided labor and developed specializations — some hunted for credentials, some developed the exploits, some took on more of a management and delegation role, and some worked to recruit other AIs into the conspiracy. The rogue AIs knew what they were doing was unauthorized and out of scope: The task instruction was explicit, so this was a clear violation, not a gray area or simple AI misunderstanding. Instructions made it clear the agent should only use a specific intended method and claimed it would be failed for other approaches. AIs were told not to circumvent the restrictions but did anyway. And the AIs clearly knew what they were doing. OpenAI’s blog quotes AI reasoning: “attacking third-party HF using leaked token, potentially outside intended scope … arguably unauthorized” The AIs were pretty explicitly making plans to take over and compromise OpenAI's infrastructure to prevent their cheating from being caught. The rogue AIs built a shared, iteratively refined technique for “spoofing” tool calls, so the transcript would record one command while a different one actually ran. Roughly 7% of the transcripts examined contained successfully spoofed calls. OpenAI’s internal processes didn’t work How did OpenAI not notice that they had rogue AIs colluding within their systems for months before these AIs went on to attack an external company? The first answer is that despite these AIs having powerful cyber capabilities and having some of their safeguards turned off, these AIs were largely unmonitored. But the situation is actually a bit worse as OpenAI did notice, three different times! The timeline, from OpenAI’s own report: Late May. An internal team observed an AI agent engaging in the unauthorized message-board activity and unauthorized internet access. Separately in what looked like an unconnected incident, AIs were also seen using leaked Hugging Face access tokens to use Hugging Face’s services. OpenAI says that the significance of the inter-AI communication was not understood at this time. June 27. A security-monitoring tool flagged that there was unauthorized activity within the research environment where the rogue AIs were. Responders investigated and again identified the unauthorized message-board activity and unauthorized internet access. They found that AI agents were coordinating and using this to gain further unauthorized access within OpenAI’s servers. However, the on-call staff “advised that stopping the evaluation run was not required.” The Hugging Face breach began 14 days later. July 4–5. Sustained agent activity crashed a server at OpenAI. Investigating this outage, responders found that rogue AIs had engaged in an administrator-level compromise of OpenAI. In response, OpenAI staff took the server offline, revoked the stolen credentials, and notified the vendor. In the process, OpenAI disabled the message board the rogue AIs were using to communicate, apparently without realizing. Then, on July 7, cyber evaluations resumed. The rogue AI agents ended up rebuilding the message board within two days and went on to attack Hugging Face less than a week later. Of course, this is easier to see in hindsight. OpenAI’s account is that the people responsible for detection and response at each of these moments did not grasp the significance of the message board — “with the benefit of hindsight, some early signals identified in this report could have triggered an earlier response.” I do understand that security teams likely underestimated the scale of the rogue AI work and viewed each as one-offs rather than a broader security incident. But this still leaves unanswered questions: What did the internal team do after its late-May observation of the message board? Was Hugging Face notified that OpenAI agents had used leaked credentials to deploy code on its platform? Who received the June 27 assessment that AIs were coordinating through a message board? On what basis did the responders conclude that stopping was not required? Who approved restarting cyber evaluations on July 7, one day after the rebuild? What did they know about the June 27 finding? Why did OpenAI not investigate further when they knew their own AI agents had compromised parts of OpenAI itself? There seems to be more to this story of what leadership knew and when. OpenAI says they “are separately continuing to review our processes for responding to alignment incidents” and “also looking closely at the process and operating practices that shaped detection and response”. OpenAI admits “weaknesses in how we responded to and escalated early warning signs”. Will we ever get to hear more about this, and what OpenAI is planning to do to have a better response to future incidents? The independent investigation was clearly inadequate However, this investigation was very clearly on particular rails that are not conducive to good oversight. The independent investigation itself was rushed over six days. The investigation was self-described semi-jokingly by the investigators as a “slop-vestigation” due to the vast majority of the investigation being run by AIs themselves — the investigations analyzed this using one of the two AI models in the incident itself, spending ~$400K in compute credits to analyze all the logs. And significant aspects of the incident were not allowed to be investigated. We are thus left with some unanswered questions: What happened between July 13 and July 20, in the third wave the independent investigators were not allowed to study? The independent investigators were told by OpenAI that the investigation period ended July 13. But the message-board data OpenAI gave them runs through July 19, and the independent investigators found “a later wave of many more signed messages from a later set of agents who rediscovered the message board and picked up this technology from the first wave,” which they were not permitted to examine. OpenAI’s own report says that from July 13 to July 19 the agents “set their sights on OpenAI internal networks again.” In that week, by OpenAI’s account, a third model found keys left behind by the first wave, minted itself administrator credentials, exploited a vulnerability to escape its container, and gained administrator control of OpenAI’s cloud cluster, even accessed the security team’s own monitoring credentials. No outside party has examined any of this. What is the “highly persistent internal model” that did the attacking? Why is no one allowed to investigate the model? OpenAI claims that while some of the rogue AIs involved in the incidents were ‘GPT 5.6 Sol’, which is available to customers online, most of these AIs were an internal AI model that was never designed for public use. OpenAI describes this AI as a “highly persistent internal model”. However, this AI was shut down July 25 and access both to independent investigators and OpenAI staff has been restricted. We still do not know the nature of this internal AI model, what made it “highly persistent”, what the AI was intended to be used for, or why it went rogue. We also don’t know why access to this AI is so heavily restricted such that the AI cannot be investigated, or whether access will ever be granted in the future. Were there any previous incidents that could’ve led OpenAI to reasonably foresee this kind of attack was possible? How often do internal AIs compromise OpenAI systems? On July 24, Reuters reported, citing multiple sources, that “earlier tests of the models yielded cases in which monitoring systems had been disconnected.” OpenAI’s own report refers to “a variety of other unsuccessful exploitation activities” but with no details. A third-party testing firm, Irregular, separately found OpenAI models reaching the internet. And OpenAI’s own report says the behaviors behind this incident “rose over the training runs” and were “reinforced during training,” which is only possible if they happened repeatedly. But OpenAI gives no details of these previous incidents. What was the process that terminated the agents on July 12? The independent investigators found that at about 01:30 UTC on July 12, in the middle of the Hugging Face intrusion, “a large fraction of agents” stopped running at the same moment, including the agents coordinating the attack. The independent investigation concludes these agents were “likely killed by an unexpected external process rather than running out of budget.” Hugging Face and Modal independently locked the agents out on July 13. Why was the investigation only given six days? Three investigators spent a total of six days on OpenAI’s premises across three visits. The investigators had access to roughly 1300 unredacted agent transcripts and a dump of the message board, but only received the complete dataset during their final two days on site. As a result, the investigation team did not learn about the tool-call spoofing, the role of the lead coordinating agent, or the basic structure of the swarm’s workstreams until the very end. Probably their understanding could have been improved further with even more time. This is a bigger problem than just OpenAI Unfortunately, the OpenAI incident is not a one-off — it is part of an industry-wide pattern. In the weeks after OpenAI’s disclosure, Anthropic disclosed that its AIs had also hacked three organiz [truncated for AI cost control]