AI News HubLIVE
站内改写5 分钟阅读

待翻译:SPAR – Fall 2026 AI Safety Research Projects

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:SPAR Research Projects 246 projects Introspection Training for Verbalization Activations Belinda Li · Anthropic We train models to be more faithful by training their verbalizations to be consistent with their internal a…

来源Hacker News AI作者: agucova

AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。

SPAR Research Projects 246 projects Introspection Training for Verbalization Activations Belinda Li · Anthropic We train models to be more faithful by training their verbalizations to be consistent with their internal activations, e.g. "thoughts" Chain of thought Mechanistic interpretability Scalable oversight Accelerating Democratic AI Leadership by Reforming the US Extraterritorial Surveillance Regime Michelle Nie · Center for a New American Security (CNAS) Keeping the most advanced AI within the control of democracies depends on integrating democratic allied nations into the US AI stack, but US extraterritorial surveillance and data access laws are eroding those allies' trust in its tech stack. This project aims to propose a reformed regime that reconciles legitimate US security interests with the trust and assurance allies need to build on American infrastructure. International governance US policy AI strategy Characterizing Attention Heads via Program Synthesis: Toward Scalable Mechanistic Interpretability David Bau · Northeastern University Can we translate a transformer's internal computations into human-readable Python programs? Building on recent program-synthesis approaches to interpretability, this project will strengthen the existing pipeline for QK circuits and tackle the open problem of grounding OV circuits in an interpretable basis, aiming for a full characterization of attention heads across open-source models. Mechanistic interpretability Economic Impacts of Frontier AI Rishi Bommasani · Stanford University Reasoning about the economic impacts of frontier AI: how does AI diffuse through work, how do new tasks emerge, how does household use substitute for labor, how do technical benchmarks relate to economic indicators. Economics of AI Societal impacts Distinguishing progress in data and algorithms Robi Rahman · MIRI Technical Governance Team We will perform dataset curation, synthetic data generation, and LLM training, fine-tuning, and evals to distinguish and quantify the effects of data improvements, separately from progress in algorithms and architectures, on increasing AI capabilities. AI strategy Compute governance Comparing Welfare Across Animal and Digital Minds Ivy Gilbert & Jeff Sebo & Bob Fischer & Toni Sims · New York University Center for Mind, Ethics, and Policy / New York University / Texas State University / Rethink Priorities / NYU Center for Mind, Ethics, and Policy This project will examine whether and how biological and artificial systems can be compared on a common welfare scale. As a CMEP-Rethink Priorities collaboration within the broader Moral Weight Project 2.0, it will assess substrate-general theories of welfare intensity, evaluate compute and related intersubstrate metrics, and map structural disanalogies that complicate comparisons between animal and artificial minds. AI welfare Philosophy of AI Extending Agentic Biosecurity Evaluations SecureBio AI · SecureBio Mentees will develop new agentic biosecurity evaluations and extend existing ones toward longer, end-to-end tasks. Biosecurity Evaluations Safe vs dangerous inference Robi Rahman · MIRI Technical Governance Team If we want to govern what types of inference are allowed, how do we define what is allowed, and enforce restrictions on what is not allowed? Technical governance Misuse risk Compute governance Measuring Headroom in Adversarial Evaluations Jamie Hayes · Google DeepMind Recent frontier system cards report that automated red teaming is saturating near 0% Attack Success Rate on jailbreaks and prompt injections, making it hard to tell whether current attacks are simply too weak or if our safety benchmarks are toy-like and eval-aware. Estimating the headroom a better attack would achieve is difficult without explicitly designing informative upper bounds. To resolve this, we will build a ladder of powerful, relaxed-constraint red-teaming attacks—ranging from continuous embedding-space PGD to internal activation steering—to quantify unexploited attack headroom and distinguish true semantic robustness from search limitations. AI security Evaluations Misuse risk An Exploration of What Kinds of Training Pressure Cause COT Obfuscation Cody Wild · Google DeepMind Study the obfuscation impact of different misbehavior monitors used as rewards, to better understand how well the emergent norm of not training against thought regions specifically matches the actual contours of obfuscation risk. Chain of thought Scalable oversight Value drift under recursive training loops Lionel Levine & Rauno Arike · Cornell University / Aether Research A controlled study of how model values change under recursive self-improvement. Alignment Behavioral evaluation of LLMs From Reading Lies to Catching Liars: On-Policy Training for Deception Probes Ann-Kathrin Dombrowski · FAR.AI Deception probes are usually trained on off-policy data — text the monitored model never generated — which is known to hurt generalization to real deceptive behavior. We test whether activation steering can fix this in two ways: by generating on-policy deceptive data, and by steering the model while it reads existing off-policy datasets to make their activations appear on-policy. Mechanistic interpretability Scalable oversight Behavioral evaluation of LLMs Can Follow-up Questions Catch Missing Reasoning? Pierre-Luc St-Charles · LawZero We have built a shortcut-following model organism that often fails because it never carries out necessary reasoning that would expose a misleading cue. This project will build a small investigator that asks targeted follow-up questions, then test whether active elicitation catches these failures more reliably and cheaply than passive judges or simple debate/consultancy. Scalable oversight AI control Chain of thought Actually Constitutional AI Seth Lazar · Johns Hopkins University This is a series of projects within the MINT lab, unified by a broad commitment to enabling liberal democratic societies to navigate the transition to powerful AI with their core values intact. This means not just (as everyone now recognises) building in some form of popular sovereignty, but also ensuring that AI systems actively work to protect and advance individuals' fundamental liberal rights. Note: these are all projects that my lab will undertake at some point; the goal is to find researchers who are interested in working on some subset of them, not to cover them all with this fellowship. Societal impacts Behavioral evaluation of LLMs Philosophy of AI Strengthening European AI Safety Unigroups Manon Kempermann · Generator Residency Analyse the current state of European AI safety uni-groups and build a playbook on how to grow and expand this talent pipeline. Generalist Differential Data for Automated AI Safety Research Alec Harris · Pivotal Research Extension Explore what data could differentially train future models to conduct useful AI safety research. Generalist Alignment Developing Updated University Group Fellowship Curricula Nixon Hanna · Generator Residency, MIT AI Alignment Survey the current curricula and develop updated versions for university groups to easily be able to pick up and run. Would likely include a general reading group (similar to BlueDot AGI Strategy) and possibly technical, governance, and/or futurism tracks. Fewer applications Generalist AI Safety Said Simply Kaustubh Kislay & Christine Corry · Wisconsin AI Safety Initiative / Second Look Research and XLab @ UChicago Explaining complex AI safety topics in short, engaging forms, including demos, writeups, and videos. Generalist Communications Revitalizing AI Lab Watch Harshul Basava & Meru Gopalan · GT AISI / GT AISI, XLab Monitor and evaluate safety practices at frontier labs by scrutinizing public-facing outputs. Fewer applications Generalist Lab governance Resources and Playbooks for University Groups James Lester & Nixon Hanna · Generator Residency / Generator Residency, MIT AI Alignment Grow and update the body of shareable knowledge and advice available to university groups (especially more comprehensive "playbooks"), so that new organisers aren't starting from scratch. Fewer applications Generalist "What's Next?": A Guide for Anyone Interested in Upskilling and Getting a Job in AI Safety Roman Ross · Generator Residency, Purdue Create a clear, efficient guide that tells people where to go next to upskill and get a job in AI safety, routing them to the programs best fit for their skill level and career trajectory. Recently added Generalist Curated AI Safety Comment Threads Helena Tran · Constellation Create an annotated collection of valuable LessWrong/EA Forum comment threads that explains the main disagreements, arguments, and cruxes in AI safety debates. Fewer applications Generalist Launch an AI Safety Comms Fellowship for Undergrads Noah Birnbaum & Sydney Von Arx · Constellation, UChicago EA / Nightingale, MATS Launch a scrappy comms fellowship for undergrads: raise the funding, recruit comms-inclined students through AI safety group organizers, and run the inaugural cohort. Recently added Fewer applications Generalist Communications Identifying function-relevant signatures in protein models for biosecurity screening Isha Harris & Gary Abel · Fourth Eon Biosecurity Institute / Fourth Eon Biosecurity Institute; the Johns Hopkins Center for Health Security This project explores biological foundation models for biosecurity screening, applying interpretability methods to identify biophysically relevant features that can reinforce screening against engineered and AI-designed biological threats. Biosecurity Mechanistic interpretability Attribution Across the Biological Threat Pipeline Anemone Franz · University of Notre Dame Most discussions of genetic engineering attribution focus on post-hoc forensic identification — tracing an engineered pathogen back to its source after release. This project will map the stages between initial intent and execution of a biological attack using engineered pathogens, and examine whether and how identification, evidentiary, or accountability mechanisms could apply earlier in that pipeline, not only at the point of forensic investigation. Biosecurity Misuse risk Faithfulness, Self-Knowledge, and Introspection Noah Siegel · Google DeepMind To what extent can we trust model self-explanations? Are models able to make use of privileged self-knowledge, e.g. via introspection or metacognition, and does this have implications for model welfare? Chain of thought AI welfare Behavioral evaluation of LLMs Transcript Analysis for Cybersecurity Evals Jack Kengott · SaferAI SaferAI and a partner collaborate on a transcript-evaluation framework to assess adversarial cybersecurity capability by analyzing transcripts from evals and real-world environments. This could mean extracting some structured data from transcripts like token efficiency, tool usage, or refusal circumstances. It could also extract less structured data like action classification (i.e. lateral movement vs execution) or highlight high-risk actions that may increase the likelihood of detection. The goal is to turn transcripts into quantitative inputs for our risk models — treating transcripts as a new KRI alongside benchmarks. Cyber risks Evaluations Technical governance Stress-Testing First Amendment Barriers to AI Regulation Alex Mark · Cambridge Boston Alignment Initiative AI regulation may implicate the First Amendment. While the First Amendment protections afforded to AI models, companies, and users are uncertain, any regulatory scheme must contemplate First Amendment litigation risk before these questions reach a court. US policy National policy Measuring and Intervening on Grader Awareness Jannes Elstner · Apollo Research We will measure grader awareness of [truncated for AI cost control]