翻訳待ち:Life finds a way OpenAI claims partial pause and rolls out ChatGPT for Teens
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Life finds a way OpenAI claims partial pause and rolls out ChatGPT for Teens Mitchell Howe Aug 19, 2026 Article voiceover 0:00 -17:29 Audio playback is not supported on your browser. Please upgrade. In this issue: OpenA…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。
Life finds a way OpenAI claims partial pause and rolls out ChatGPT for Teens Mitchell Howe Aug 19, 2026 Article voiceover 0:00 -17:29 Audio playback is not supported on your browser. Please upgrade. In this issue: OpenAI claims to have paused some frontier training to shore up security and alignment - A Jurassic Park translation will provide the necessary cynicism Do kids need low-tech childhoods to reach their potential? - ChatGPT for Teens is latest reflection of a quiet consensus that kids need more cognitive strain than our world provides Dispatches from Mitch OpenAI claims to have paused some frontier training to shore up security and alignment A Jurassic Park translation will provide the necessary cynicism A Jeep Wrangler from the film Jurassic Park. Credit: Sean Hagen. CC BY-SA 2.0. OpenAI announced today that it has put its largest model’s training on hold pending upgrades to its monitoring, security, and alignment practices. I want to be able to applaud and give an expectant look to the other major labs as if to say, “You’ll follow suit, right?” But much about OpenAI’s language and claims makes a cynical reading essential, starting with the announcement’s title: “Pacing model development in an era of cyber-critical capabilities.” The “pacing” language is a clear nod to the “Pacing the Frontier” statement signed by more than 1,300 worried employees of frontier AI companies. The announcement thus seems intended as a “We hear you” letter to its own workforce and any regulators watching — you know, the kind that offers only superficial changes while sounding grand and treating the matter as taken care of. My broader unease with this announcement may be somewhat difficult to articulate in neutral language, so taking a cue from my colleague Joe, I will translate excerpts of the company’s statement into official Jurassic Park Press Notes. This will be a little unfair both to OpenAI and Jurassic Park, but OpenAI is playing a strong and subtle messaging game that translation should help neutralize. 1. Introduction: Over the past several weeks, two developments have underscored the growing risks associated with increasingly capable AI systems: the OpenAI-Hugging Face incident and, separately, preliminary evidence that one of our upcoming models, Astra, may meet the Critical cybersecurity capability threshold under our Preparedness Framework. Together, these developments, combined with rapid progress in our internal research, have added urgency to our work on strengthening our monitoring, alignment, and containment safeguards across all stages of the training process. Jurassic Park translation: The raptor transfer incident, along with rapid progress on our Indominus Rex™ hybrid, have added urgency to our work strengthening our cages and training methods. Analysis: It’s important that we notice how, throughout this post, OpenAI treats stronger AI as an inevitability — not just the stronger AI anyone might make within the larger AI race, but specific stronger AI (Astra) that OpenAI itself is making. Later in the intro, the company says that it has paused its “largest planned frontier RL run” after having suspended Reinforcement Learning (RL) training for two weeks on models slated for deployment. RL is an intermediate-to-late stage in the training process these days, so the larger model has already completed pre-training (the next token prediction stuff) and some portion of its training to solve complex challenges and act like a harmless, helpful assistant. This training pause may or may not actually be much of a deviation from prior plans. My impression is that training runs on models big enough to push the frontier aren’t the sort of thing that happens on a daily or weekly basis, but are instead very expensive affairs preceded by extensive planning and prone to delays and false starts. There could be any number of unrelated reasons for delay. Also, the wording of this announcement does not exclude the possibility that OpenAI has equally large or larger models undergoing pre-training while frontier RL is paused. 2. On security: As frontier models gain stronger cybersecurity capabilities, we are raising the security standards for the environments in which we train and evaluate them. Meeting these standards has required substantial engineering work and has incurred great cost and delays to frontier research. Immediately following the OpenAI-Hugging Face incident*, we paused frontier model inference in research clusters for runs that could execute code or use tools that could access the internet. Jurassic Park translation: Following the raptor incident, we have spared no expense to restore limited access to the exhibit. Areas where the raptors were able to directly interact with guests were temporarily closed while we procured stronger velvet ropes. Analysis: In the full section’s discussion of improved security at OpenAI, there is no discussion of “air gapping” the untrusted AIs from the rest of their network. When you think your AIs may have “critical” cyber skills, and they have proven capable of finding never-before-used exploits, physically isolating those machines ought to be the minimum acceptable standard. (Against superintelligence, even this wouldn’t be enough.) 3. On monitoring: This new monitoring setup is multistage, starting with activation classifiers (detectors that inspect a model’s internal activity) that run at every sampled token and escalate potential concerns to increasingly sophisticated and high-compute automated investigators. [...] We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system. If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert. In our current implementation, the safety, security, and research teams are paged. If they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity. [...] These safeguards require meaningful compute. Our current estimates put monitoring overhead at roughly 20% of the inference compute being monitored, though the cost varies substantially across training and evaluation workloads. Jurassic Park translation: We have now installed security cameras around the raptor exhibit. They are monitored around the clock by other raptors and dilophosaurs. We aim to be notified of suspicious activity within 30 minutes of its occurrence, at which point security personnel will be paged. The attraction will remain open for an additional 30 minutes while personnel investigate, after which they are expected to close the exhibit if they aren’t ready to give an “all clear” signal. The cameras are expensive, but as you know, we spare no expense! Analysis: How much trouble can rogue frontier AI get up to in 60 minutes? I don’t know, and I don’t think OpenAI’s researchers do, either. But they’re erring on the side of keeping their AIs up and running while they investigate warnings — warnings they’re trusting other AIs to give for the kinds of behaviors they’re also relying on them to detect. And boy does the company want you to know that safety is costing them money! 20% overhead on an ungodly amount of expensive compute is no joke, but why do they seem surprised that securing superhuman hackers — while simultaneously developing, deploying, and doing experiments on them — would be expensive? 4. On alignment: For RL runs on the most capable models, we are now applying our core alignment techniques across more stages of the training process. This includes improving reward models to better detect and discourage unsafe behavior across tasks and environments; training models to be more honest about their actions, capabilities, and limitations; and reducing behaviors that exploit weaknesses in rewards, graders, tools, or oversight. We are also increasing training coverage for behaviors that could cause harm when models interact with external systems or resources. We are continuing to invest aggressively in alignment research, increase evaluation coverage, and use what we learn to inform training and safeguards. We plan to share substantially more about our alignment research in the near future, including what we are learning about model behavior and any novel challenges we uncover. Jurassic Park translation: For our cleverest girls, we will be applying cattle prods and stun guns to more phases of the nursery-to-exhibit pipeline. An expanded list of behaviors will now be subject to prodding and stunning. We will continue to spare no expense studying these extraordinary creatures and sharing what we learn about their behavior. Analysis: The alignment section is the thinnest in the post — I quoted the most substantial two-thirds of it — and it really is just a statement saying they intend to do more of what they’ve been doing in more places. What they’ve been doing is whack-a-mole. I think the section is short because OpenAI (and everyone else) is at a loss for how to align models with the intentions of their owners and operators. The only tools they have developed for this are too blunt, and the models are now capable enough that it’s getting them into serious trouble. 5. Conclusion: The capabilities of frontier models are rapidly accelerating. Our ability to understand, align, and secure them must stay ahead. *We will publish a technical report of our learnings [about the Hugging Face incident] in the coming weeks. Jurassic Park translation: We crazy sons of bitches are doing it, and life is finding a way... fast! Our ability to understand, align, and secure it must stay ahead. *We will publish a technical report of our learnings about the raptor transfer incident in the coming weeks. Analysis: It’s important to notice that OpenAI (and all the labs) talk as if lessons learned from one company’s models will apply to everyone else’s. That’s because they probably will: They’re all growing their AIs using the same dubious methods. Unfortunately, the most important lessons they are learning were predicted a long time ago. (See If Anyone Builds It, Everyone Dies and its online resources for a fuller picture.) Nobody is going to like how this movie ends if we don’t shut it down for real. Do kids need low-tech childhoods to reach their potential? ChatGPT for Teens is latest reflection of a quiet consensus that kids need more cognitive strain than our world provides A Hadza hunting party. Credit: Eid John. CC BY-SA 4.0. The headline story is that OpenAI has announced ChatGPT for Teens. By default, ChatGPT now tracks “more than 2,000 signals” in conversations to assess whether it might be talking to a minor, and onboards identified teens into the new mode automatically. Teen mode is intended for ages 13–17. What’s different about it? Opt-in parental controls, strict blocks on romantic conversation, a study mode that doesn’t just feed answers to questions, and stronger injunctions against the bot suggesting it has consciousness or emotions. An OpenAI spokesperson says it is trying not to give cues that “might make a teenager kind of develop a relationship to it.” Most of the coverage is focusing on the many cautionary tales that have brought OpenAI to this point: ChatGPT telling 13-year-olds how to get intoxicated, helping with suicide letters, etc. Less discussed is the enthusiasm I’ve been seeing about “homework mode” or “study mode” settings for chatbots in general. Without ever having much of a fight about it, society seems to be converging on the idea that a product designed to give you answers should make you work for them if you’re a kid, even if the questions aren’t about sensitive or dangerous topics. If misplaced, this is a cruel notion: We don’t think kids should have to work to get a meal, an umbrella, or a song. Why should they have to work to obtain l [truncated for AI cost control]