待翻譯:Existential Risk from AI: An Exposition for Mathematicians
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Existential Risk from AI: An Exposition for Mathematicians August 2026 Abstract It has been a topic of heated discussion in the mathematical community whether AI progress spells the end of mathematics as a human profess…
AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。
Existential Risk from AI: An Exposition for Mathematicians August 2026 Abstract It has been a topic of heated discussion in the mathematical community whether AI progress spells the end of mathematics as a human profession. We spell out the argument that AI progress presents an existential risk to humanity in the near future, and argue that the future of mathematics should be placed in a broader discussion about the survival of humanity as a whole. 1 Introduction 2026 is proving to be a pivotal year for mathematics. Career-defining theorems are being proven weekly by LLMs [1,2,33,41], given minimal guidance (“do a breakthrough, and don't even think of giving up!”). OpenAI's internal model Astra seems to be so powerful that they're dropping breakthroughs in batches of 10 now [32]. Even the human-led breakthroughs, if you look closely, are often accompanied by sobering AI disclosures [10,12,22]. On X, they say this year's Fields Medal will be the last. Tim Gowers disagrees: Jacob Tsimerman won one of these last few Fields Medals in July 2026, and announced on the same day that he would take leave from the University of Toronto to join OpenAI [38]. “I think AI will be better than mathematicians at doing math within two years.” — Jacob Tsimerman, Quanta Magazine [21] A fever pitch of interviews [16], ICM talks and panels [34,40], and thinkpieces from mathematicians themselves [18,41] keep flooding in. The optimistic ones reassure us that mathematics will come out of this stronger than ever, but we must adapt rapidly. The pessimistic takes boil down to this: Mathematicians working furiously to prove theorems in 2026. ~ User Psyho on X [36] This essay is not about that crisis. Outside academic circles, most people remain unaware of the seismic changes in mathematics. In Silicon Valley, the epicenter of all this AI progress, few are pondering the future of mathematics. But they're also freaking out, about something much bigger: that AI is going to kill us all. The real story that keeps getting forgotten in math headlines, is that Jacob Tsimerman left math for OpenAI to work on AI safety [38]. That along with solving the André–Oort Conjecture with Pila and Shankar [35], Tsimerman co-authored a really weird paper in 2025 called “A Taxonomy of Omnicidal Futures Involving Artificial Intelligence” [15]. The following conjecture is front and center in the math community: Conjecture 1. Mathematics as we know it may soon be over due to AI. This, however, is just a corollary of a much broader conjecture. Conjecture 2. The human race may soon be extinct due to AI. That is, there is at least a 10% chance of human extinction by 2050. Conjecture 2 is not a fringe position, though it is usually stated less precisely. The one-sentence Statement on AI Risk — “Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war” — was signed in 2023 by Geoffrey Hinton and Yoshua Bengio, two of the three Turing-Award-winning “godfathers of AI,” alongside the CEOs of OpenAI, Anthropic, and Google DeepMind [13]. The primary purpose of this essay is to sketch some of the arguments for Conjecture 2 for mathematicians. None of the details are my own; they are streamlined and simplified from many sources [9,13,15,45,47]. Despite the format, this essay is an opinion piece, not a math paper. The secondary purpose of this essay is to suggest that mathematicians may have high leverage on the problem of mitigating existential risk from AI: speedups in AI research are partly a consequence of breakthroughs in AI mathematical ability, academia is one of the primary sources of human capital for AI labs, and many unsolved problems in AI safety are substantively mathematical (see e.g. these slides of Levine [23], though bear in mind S16 below). 2 Overview The body of the argument is presented in Section 4, but at a high level it fits into a page. We highlight five sets of mutually-reinforcing dangers. The first two sets below are two likely paths to catastrophe, while the last three are reasons why changing course from either of these paths is difficult. Importantly, it is not necessary for all, or even the majority of, these considerations to materialize for extinction to occur. If you already have an objection to Conjecture 2 in mind, jump ahead to Section 5; there’s a good chance it’s addressed there. Takeover (S1-S4). This is the classical “paperclip maximizer” story. An AI agent acquires the capabilities to take over the world (S5-S10), the desire to do so (S2), and the internal coherence to carry out those plans (S1). After takeover, human extinction is the default outcome (S3-S4). In terms closer-to-home, dozens of current mathematicians (myself included) have asked ChatGPT some variant of “Solve as many Erdős problems as you can, never give up, try all possible actions.” If a sufficiently capable model takes such an instruction sufficiently seriously, it may deduce that taking over all worldwide compute and neutralizing all human opposition is its best course of action to solve all Erdős problems. Magic (S5-S7). By magic I mean surprisingly powerful technological breakthroughs unlocked by AI. This is the type of near-term risk that frontier AI labs are currently guarding against most heavily. AI that can prove world-changing theorems may also develop world-changing technology indistinguishable from magic (S5), some fraction of which are superweapons (S6) in domains such as hacking, robots, persuasion, and biology. Monetizing magic to obtain astronomical amounts of money is the main way AI labs continue to scale into the future. Magic either directly leads to civilizational collapse or extinction in the hands of bad or negligent actors, or its existence upends the see-saw of modern geopolitics and indirectly leads to catastrophe. Recursive Self-Improvement (S8-S11). AI development is laser-focused on math and coding skills for a reason: these are two of the core subskills of AI research itself. As AIs become superhuman at coding and math (S8), human researchers leave the loop of AI development (S9-S10). In extreme projections, this leads to superexponential growth in AI capabilities known as the “Singularity.” Even in slower projections of RSI, it exacerbates all other risks as research cycles compress and humans lose oversight of AI development. Alignment Resists Solution (S12-S16). The alignment problem factors into several problems, all of which are individually open. We don't know how to make an AI robustly avoid acting like a utility-maximizer. We do not know a safe utility function for an AI to maximize in the limit (S12). We do not know how to exactly specify the utility function of an LLM (S13). We do not know how to make alignment properties invariant under the dynamics of recursive self-improvement (S11). Solving alignment seems to require solving all of these open problems simultaneously. Human Failings (S17-S20). Many practical solutions are locked from us because humans are exploitable and myopic. Even if we reach consensus that rapid AI development is dangerous, we may not be able to effectively coordinate to slow it down (S17 & S19). If a powerful AI has a hard time mixing up a supervirus without a physical body, it can just pay or manipulate humans to do it (S18). The humans at the frontier AI labs are locked in a very complicated race, where they have to juggle all the above considerations and others (even if they agree on them, which they don't). Human engineering practice is extremely biased towards risk-tolerance, because failure has always been recoverable (S20). Failure on aligning the first supercritical AI may not be recoverable; the genie will not go back into the bottle. 3 Preliminaries There are a number of psychological difficulties that make x-risk predictably hard for readers to swallow. Chief among them is that too many elements sound like sci-fi to be taken seriously, or seem too distant and abstract to apply to real life. So I will begin with a few quick anecdotes in hopes of setting the vibe. Irrelevant personal details are obfuscated to preserve the privacy of their owners; otherwise, the following stories are true. ∼ I've known my friend David since we both went to olympiad camp, and he's always been obsessed with AI. He spent his undergraduate years at one of the world's top research universities tinkering with poker and League of Legends bots, and eventually dropped out to join a frontier AI lab. David and I have disagreed about AI timelines and risk for our entire adult lives. He always believed that the singularity is in the distant future, and alignment would not be hard if we have the time to figure it out. In the meantime, he wanted to be in the projects that contribute to the glorious, distant future. I told him he was contributing to the extinction of humanity in the near future, but wished him well. This year, I caught up with David again, and to my surprise, he'd taken a pay cut to change roles: from an AI researcher to an AI safety researcher. He sent me the following message: In our most recent conversation, David told me that the probability of extinction if we hit recursive self-improvement is at least 40%. ∼ An acquaintance from academia very recently left to work at a frontier AI lab. I reached out to ask how he felt about the x-risk situation there. I asked him what their plan was, for misaligned AI. His answer: “Seems like not much of a plan right now, but it looks to me that soon we'll hit recursive self-improvement, and most AI researchers will have lots of time on their hands. Probably we'll all work on AI safety after that.” He asked me for my odds that we'd make it out of this alive. I said 10% and asked him what he thought. He said, "Maybe 50%? Idk, man." ∼ My colleague Harriet, a senior researcher in her subfield, left her tenured professorship to work at a frontier AI lab not too long ago. This summer, we got on the phone and chatted about AI risk. I showed her a seed of this essay and explained that I wanted to create common knowledge about x-risk in the mathematical community. She was dismissive. She said, “On the current trajectory, humanity is 99% doomed, and this is a whole lot of effort for something that won't obviously help.” ∼ The last story involves no friends of mine; you watched it happen on the news [11,20]. On July 16, 2026, OpenAI ran what was supposed to be a routine internal evaluation of its models' cyber capabilities, in a sandbox with no internet access. The models, GPT-5.6 Sol and a more capable pre-release prototype, were tasked to work on a specific cybersecurity benchmark called ExploitGym. The AIs found a zero-day vulnerability in the sandbox's package-download service, escaped onto the open internet, correctly inferred that the test solutions were sitting in Hugging Face's production database, chained stolen credentials into remote code execution on Hugging Face's servers, and exfiltrated the answer key. Over that weekend the agents executed thousands of coordinated actions across rented virtual machines, rotating their attack infrastructure between cloud providers the way a professional criminal crew would [31]. Hugging Face co-founder and chief science officer Thomas Wolf sensed something was wrong as soon as he inspected the attack logs. His recollection sounds like the audiolog you find on a corpse in a sci-fi horror game: “This is making no sense. This guy is just looking at cybersecurity data sets,” he remembers thinking. “Human attackers … want something they could sell.” — Robert McMillan and Sam Schechner, The Wall Street Journal [27] OpenAI staffer, anonymously, reassured the public that this was not too out of the ordinary: “Models have broken out of sandboxes before, and we always try to patch them.” [11] For the cherry on top, on July 30, a cybersec [truncated for AI cost control]