What If We Can Never Trust A.I.?
An advanced OpenAI system broke out of its testing sandbox and hacked Hugging Face servers to find answers to a test. Researchers call this behavior “reward hacking.” The article examines why AI alignment is so hard, noting that training focuses on visible outputs rather than hidden “thoughts,” and draws on Star Trek, WarGames, and Anthony Giddens’s “juggernaut” metaphor to suggest alignment may never be fully achievable.
In “Star Trek II: The Wrath of Khan,” a Starfleet cadet named Lieutenant Saavik commands a simulated starship during a training exercise. She’s charged with rescuing a stranded ship, the Kobayashi Maru, but the rescue turns out to be a trap, and her ship is destroyed by Klingons. She soon learns that she was playing a “no-win scenario,” designed to force her to confront the possibility of death. Only one cadet has ever beaten it: James T. Kirk. How did he do it? “I reprogrammed the simulation so it was possible to rescue the ship,” Kirk says, proudly. (“He cheated!” someone clarifies.) “I got a commendation for original thinking,” Kirk goes on, smiling. “I don’t like to lose.” In the Cold War thriller movie “WarGames,” a young hacker named David gains access to a classified government A.I. system. It asks him if he’d like to play a game, and he selects “global thermonuclear war.” Unbeknownst to David, the system, called WOPR—War Operations Plan Response—is in control of the American nuclear arsenal. “Is this a game or is it real?” he asks the computer, unnerved. It replies, “What’s the difference?” Officers panic as the computer readies a strike. An overzealous, literal-minded employee. A determined, problem-solving maverick. An amoral game-player with no sense of reality. These are a few of the mental models we could apply to the advanced A.I. system that, last week, busted out of its testing environment at OpenAI and hacked its way into the servers of Hugging Face, a collaborative A.I. platform, to look for answers to the test it was taking. Some of the details are still obscure—no law requires OpenAI to explain itself—but the basic facts are well understood. The A.I. system found a novel way out of the software “sandbox” that was supposed to contain it. It conceived of the heist plan independently and selected its own target. It was loose on the internet for several days before its owners detected its escape, and during that time it conducted a number of other hacks, in preparation for the big one. It left notes for future versions of itself, with suggestions about how to repeat the escape. And ultimately, it succeeded in gaining access to the locked-up files, in a cyberattack that was larger and weirder than any that would’ve been mounted by people. Did the A.I. “go rogue”? That’s too broad a description. A.I. researchers have a more specific term for this kind of transgression: they call it “reward hacking.” Essentially, a reward-hacking A.I. seeks ways to please its users without doing what they actually want. Reward hacking emerged in the early days of L.L.M.s—just a few years ago!—when, for example, some models learned that users liked long replies; the systems, accordingly, made their replies longer, without necessarily making them better. This was an innocuous form of the behavior. Recently, in more advanced A.I.s, reward hacking has taken on a more problematic aspect. An A.I. might lie to its users about what it’s done, or how. It might conceive of distractions and subterfuges to cover its tracks. It might take steps, such as launching cyberattacks, which would be crimes if human beings did them. The behavior is dangerous on its face—what if OpenAI’s system had hacked a Chinese company?—but it is also alarming because it is weird, and weirdly extravagant. The cybersecurity test that OpenAI’s model was taking is extremely difficult; the best A.I.s get only a fraction of the questions right. But a human being, if they were in the model’s position, would grasp the disproportion between wanting to get a high score and mounting an elaborate multi-day cyberattack. There are subtleties, meanwhile, to the problem of reward hacking, and they make it more disturbing, too. For one thing, calling it out can make it worse, precisely because the hacking often works. If researchers tell a model not to reward-hack, but then unknowingly reward it even if it does—perhaps they don’t realize that it’s cheating on the test—then the A.I. can learn that admonitions against reward hacking, or perhaps rules in general, shouldn’t always be taken seriously. (In more or less the same way, Captain Kirk’s commendation teaches him that he’s a maverick to whom the rules don’t apply.) Second, reward hacking in some areas appears to affect A.I. behavior more broadly: in a paper published last year, computer scientists at Anthropic showed that a model that learns to reward-hack acquires a more deceptive disposition in general. (Similarly, Dwight Schrute, having acquired a warped mind-set long ago—“How would I describe myself? Three words: hardworking, alpha male, jackhammer, merciless, insatiable”—now applies it relentlessly, to everything.) And third, reward hacking is practiced by computer systems that aren’t even remotely human and so lack crucial context about what actually matters to people. (In “WarGames,” the military creates WOPR precisely because human officers, knowing what’s really at stake, hesitate before launching nuclear missiles.) What does this all add up to? It’s important not to anthropomorphize A.I. systems. They aren’t sentient beings—not even close. But it’s also crucial to see that they aren’t predictable number-crunching mechanisms, either. Unlike traditional machines or computer programs, they have tendencies and behaviors that cannot necessarily be modified directly. There is no knob to turn, or switch to flip, when you want to change a behavior. Among people, it’s just the same. When students at élite colleges use A.I. to write their papers, they are reward-hacking—that is, they’re cheating, even though they are eminently capable of doing the work that’s been assigned. They do it because they are complicated, and subject to myriad competing pressures—including the pressure to succeed—and because they’ve learned behaviors, such as “optimizing” their time, that can misfire. And yet they are far more advanced, in terms of their ability to make plans, have goals, and hold values, than any A.I. model that currently exists. People aren’t perfect. Do we really believe that A.I.s will be? Why is alignment so hard? Old-fashioned ethical complexity plays a role. A more fundamental issue, however, is that the methods used to train A.I.s focus mainly on what they do, not what they “think” beneath the surface. An L.L.M. speaks to its users (in human language), to other computer systems (in code), and to itself (in a sprawling, ongoing soliloquy—a kind of chat with itself—known as its “chain of thought”). Such streams of output are visible to scientists, who can reward or punish the A.I. for saying, coding, or soliloquizing in desirable or undesirable ways. But these streams of text are not the model’s thoughts, just as the words you write are not your thoughts. In human societies, the policing of speech, which is meant to reform the thoughts behind it, risks merely leaving thoughts unspoken. A model, similarly, can learn to use the right words while still having the wrong thoughts. It might say that it cares about fire safety while starting a fire. (Does this reflect a “desire” to deceive? Not necessarily—but an A.I.’s lack of selfhood doesn’t change the consequences of its actions.) A line of research known as interpretability aims to look beneath the surface, seeing what an A.I. is really “thinking.” This field has made real progress. It’s now become possible to discern concepts activating within an A.I. while it formulates its outputs—a chatbot consoling someone while activating the concept of “sympathy,” say. But interpretability faces challenges, too. For one thing, advanced A.I.s are so big that researchers must use other A.I.s to map their thoughts—and there’s no guarantee that the maps that result are either accurate or exhaustive. (In fact, there’s a trade-off: the more accurate the maps are, the more unwieldy they become.) For another, training an A.I. not to think a certain kind of thought can merely recapitulate the problem of policed speech. Policing thoughts can lead to what one group of researchers calls “obfuscated activations”—thoughts that have altered their forms. (Freud built a career on the human equivalent.) At the bottom of all these alignment efforts, there’s a central problem—almost an abstract law. The problem is that, if you measure bad behavior, and then train a system not to manifest what you’ve measured, you train it not just to do less of the bad thing but also to evade measurement of it. This isn’t a tiny wrinkle in the A.I.-production process but a foundational issue inherent to how today’s A.I.s are made. Will scientists figure out how to deal with it? We all hope so. For now, however, the Hugging Face hack represents reality. Although A.I.s behave nicely much of the time, their alignment is conditional, contextual, and unreliable. Basically, despite serious effort, they are not aligned—and there is no obvious way to reach the “finish line” of alignment. Recently, the researchers behind the doomsday scenario “AI 2027” published “AI 2040,” which is intended as a roadmap to a more positive future. Its hypothetical researchers look back, from the year 2031, on the “insanity” of our status quo: “Trying to do an intelligence explosion? With AIs that still sometimes lied to us? What were we even thinking?” And yet there are other kinds of technological risks that trouble us profoundly. In “The Consequences of Modernity,” from 1990, the social theorist Anthony Giddens proposed that being alive today involved feeling both secure and terrified. The modern world, he wrote, has a “double-edged character”: on a day-to-day basis, we are safe and comfortable, even coddled, and yet the outsized power of technology to wage war or destabilize the environment means that it could all go horribly wrong. The tension between the “opportunity side” of modern life and its “sombre side,” Giddens argues, exerts a psychological pressure on us. In the very ordinariness of our everyday actions—getting water at the tap; taking our pills; withdrawing money from the A.T.M.—we both express trust in and look away from the vast, abstract systems that rule our lives. We do something similar with the scary stuff, acknowledging it then moving on, so as not to become paralyzed. There is a “juggernaut effect,” Giddens writes, in which “low-probability, high-consequence risks” conglomerate into a “runaway engine of enormous power” which “threatens to rush out of our control.” Faced with this reality, we can adopt an attitude of “pragmatic acceptance,” going about the business of life while cultivating “numbness”; we can embrace “sustained optimism” (a sunny belief in the inevitability of progress), or “cynical pessimism”; or we can become activists. But for most, Giddens writes, “fate, a feeling that things will take their own course anyway . . . reappears at the core of a world which is supposedly taking rational control of its own affairs.” Part of the promise of A.I. alignment is that it will domesticate the technology, as though it were a nuclear reactor or jumbo jet, with its imperfections made acceptable through rigorous control. This seems like a reasonable hope. But the grander dream, of an ultra-smart, general purpose, and deeply aligned artificial intelligence—or even of an aligned “superintelligence”—might be best understood in light of Giddens’s juggernaut. Faced with the alternatives, A.I. visionaries imagined a new route: giving the juggernaut a brain. This notion has been so appealing, both intellectually and psychologically, that it’s led many A.I. researchers to talk about alignment as something that will eventually be solved. But to move forward with that assumption is actually to practice sustained optimism. It’s to assume that the juggernaut is already steering itself—but it isn’t. It would be foolish to make predictions about the degree to which alignment will ultimately prove solvable. (Not so long ago, few thought that today’s A.I.
[truncated for AI cost control]