AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。
This essay is over 13,000 words long and represents our most substantial writing on AI safety since the original essay. Over the last few months, loss-of-control incidents at OpenAI and Anthropic have intensified concerns about AI safety. Warnings about existential risk have increasingly reached the broader public, alongside calls to slow AI development. Dario Amodei’s call to “pace the frontier” reflects the concern that safety efforts are not keeping up with AI capabilities. The most prominent example was the OpenAI - Hugging Face incident, where hundreds of OpenAI agents got access to the internet and hacked Hugging Face to find how they were being graded on an evaluation. Over the last two weeks, the details of many other instances of such activity from OpenAI agents being evaluated have surfaced, such as agents using an old Wiki website to communicate with each other despite restrictions on such activity, or attacking a software repository to attempt to upload malicious software. The AI safety community has viewed these incidents largely as a crisis for alignment.1 In this view, alignment will become harder over time as agents learn to reason covertly, and as a result, loss-of-control incidents are likely to become much more widespread and damaging as agents become more capable. On the other hand, cybersecurity practitioners have largely viewed these incidents as consequences of companies failing to adopt basic security precautions. They do not see the incidents as a sign of AI reaching a new milestone in cybersecurity. This is also the dominant reaction in the tech community outside AI, which has largely treated the incidents as a result of ineptitude and negligence by AI companies.2 Both communities have made important points. But the polarization between them is counterproductive, and there are lessons from both the alignment failure and the security failure for understanding the path forward. In this essay, we apply the AI as Normal Technology framework to synthesize the views of the safety and cybersecurity communities and offer a constructive middle ground. We think AI companies should be liable for what their agents do, and this should be clarified through policymaking. At the same time, recognizing their responsibility does not mean that preventing future incidents is a solved problem. We identify three areas where investment is necessary to address loss-of-control risks: research to develop better methods for controlling increasingly capable agents, translating existing research and known control techniques into usable tools, and organizational changes to ensure that these tools are actually adopted. In our view, standards for organizational governance should be a key way to pace the frontier and pull AI companies out of the “move fast and break things” attitude they currently operate in. This essay has three parts. In Part 1, we argue that alignment alone is not enough to prevent such incidents, and discuss technical, organizational, and policy interventions for improving AI control. We are cautiously optimistic that the right investments and policy interventions can allow AI control to keep pace with AI capability improvements. In Part 2, we discuss the impact of improving AI capabilities, such as agent swarms, on cybersecurity. In Part 3, we share how our views on AI safety have changed in light of new evidence. A summary of our argument: We agree with security practitioners that OpenAI did not take adequate protections for controlling their agents. But this is not just a matter of applying 30-year-old security methods to a new domain. Security for AI agents — AI control — while important, is not a solved problem. While known control methods would have prevented the Hugging Face incident, as agent capabilities continue to advance, we will only be able to control them if we invest adequately in control interventions. We also agree with security practitioners’ implicit position that these incidents are primarily a security story. In the AI safety community, rogue agents are treated as inherently catastrophic because of the assumption that there is an endless list of risks that will arise from their development. We disagree. We have long advocated that the best approach to AI safety is to identify the risks and address those specific risks. Over the last few months, it has become clear that one urgent risk is cyberoffense, because it has unique properties that allow agents to carry it out autonomously. We should similarly invest in defenses against other specific risks, such as biorisk and risks from military AI. We agree with the safety community that there is an urgent need for technical and policy interventions to prevent loss-of-control incidents. But in our view, marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques. More broadly, there are many common-sense policy proposals that could help promote investments in AI control where we share common ground with the safety community. How should we reason about AI’s impact on cybersecurity? It’s plausible that advances in agent capabilities upset the offense-defense balance for cybersecurity. We cannot yet be certain, but there is enough evidence that agent capabilities might soon make widespread cyberoffense possible that urgent action is warranted. We discuss potential interventions for tilting the offense-defense balance towards defenders. How the AI as Normal Technology framework has evolved over the last year. We take stock of AI progress and share how we have updated our views. In the essay, we did not pay sufficient attention to safety risks that arise during development and evaluation (as opposed to the widespread deployment of models). We were too confident that companies would take basic control precautions and underplayed the importance of jaggedness, which led us to underestimate how quickly capabilities could improve in domains such as cybersecurity. At the same time, many distinctive claims of AI as Normal Technology have held up, and it remains valuable for understanding AI’s societal impacts. In particular, we think recent incidents support our continuity hypothesis — the behavior of “rogue” agents became apparent and widely publicized while they are still far from causing serious harm and incompetent at hiding their traces. The societal reaction to even the relatively small harms from these incidents has been fierce (and the safety community deserves credit for keeping up pressure on companies). Whether this translates into meaningful changes in companies’ behavior remains an open question, and a test of the usefulness of the AINT framework. Finally, we think AINT is particularly valuable for analyzing these incidents because it provides a framework for synthesizing the AI safety and cybersecurity communities’ views into a coherent plan of action: hold companies responsible, invest in control, and strengthen defenses against specific risks. This table summarizes our argument. While the AI community has in principle emphasized the notion of “defense in depth”, in practice, alignment has by far been the intervention that received the most attention and investment. On the other hand, the cybersecurity community has viewed control as largely a solved problem. In our view, to protect against loss-of-control incidents from legitimate parties, we need to urgently invest in control so that it keeps pace with capability improvements. Neither alignment nor control help against malicious users, so we need downstream defenses and resilience to alleviate the impact and severity of risks. All of these efforts can be shored up by policy interventions. The table emphasizes interventions for cybersecurity, but we should simultaneously invest in defenses against other risks, such as biorisk. Table of contents Part 1: We can get better at AI control through technical and policy interventions Alignment is helpful but not sufficient for preventing safety incidents OpenAI did not use known control interventions that would have prevented the incident Existing organizational governance norms would have prevented the incident Why has the AI community under-invested in control? Will control keep pace with AI capabilities? AI control should become a job (and a part of every job), just like cybersecurity AI policy can incentivize AI control Part 2: There is great uncertainty about how AI will impact cyberrisk. But what we need to do regardless is relatively clear. The specific threat that is urgent is cyberrisk How will autonomous cyber capabilities and open-weight models affect the attacker/defender balance? To understand what open-weight models mean for the future of cybercriminal activity, we need to understand the past correctly The majority of cybercriminals are surprisingly low-tech, and the barrier is monetization, not exploitation So, how much will AI help cybercriminals? Are we about to see harmful autonomous agents? Why the Morris worm is a good analogy for what might be about to happen What the future of cyberdefense might look like: the roles of people and defensive AI What are the gaps in the state of cyberdefense against AI? Part 3: What AI as Normal Technology got wrong — and right — about AI safety Do risks arise from development or deployment? How should we deal with rogue agents? The continuity hypothesis Jaggedness and risk-specific defenses Is AI safety on track? Conclusion Part 1: We can get better at AI control through technical and policy interventions In this part, we focus on AI agents that are operated by legitimate actors — such as consumers, businesses, and AI developers — who don’t intentionally use them to cause harm. This was the case for the OpenAI - Hugging Face incident. In Part 2, we focus on malicious actors who want to use agents for intentionally causing harm, such as using them for cyberoffense. How can we prevent AI agents from taking harmful actions? There are two broad kinds of interventions. One is AI alignment. This involves changing the AI system itself, such as by fine tuning or using reinforcement learning with human feedback (RLHF), to make it less likely to take harmful actions or give harmful responses. Alignment has been pivotal for the commercial success of AI so far. The other intervention is AI control: interventions made outside the model weights to prevent harmful actions — even if the agent is misaligned. This includes improvements to sandbox security (to prevent AI models from taking unanticipated actions outside the sandbox), implementing the principle of least privilege, comprehensive logging, automated tripwires for unsafe behaviors, rapid shutdown mechanisms, and monitoring the agent to detect and prevent harmful actions. These mechanisms allow us to prevent unsafe actions even if agents are misaligned and try taking harmful actions. Roughly speaking, AI control can be thought of as cybersecurity against AI agent adversaries. Whereas cybersecurity has traditionally been concerned with human actors who seek to compromise or exploit a system, AI control uses these same principles to prevent unwanted actions by AI agents. Unlike cybersecurity, control interventions target agents being used within an organization (as opposed to human adversaries or external agents). This makes the problem both more and less challenging than traditional cybersecurity. It is less challenging because the agent is directly controlled by the organization (rather than being an unknown adversary), so its operating conditions can be closely monitored and intervened on. It is more challenging because these agents are often deployed by users who might have escalated privileges, and imposing security constraints also imposes constr [truncated for AI cost control]