Skip to content
AI News HubLIVE
Public articles 12Collected articles 12Trust 86Refresh 120 min
Health Auto-pausedSource type ResearchFull-text rights Official full textLast ingested 2026-08-25ID eleutherai-blogStatus Not enabled

Official research collective blog; confirm reuse terms before full body display.

Latest public articles

A Dynamical Model of AI Governability

A toy dynamical model of whether the AI workforce that builds future AI ends up cooperative or uncooperative: where the basin boundary lies, what current evidence says about which side we are on, and what would tell us we are on the good path.

EleutherAI BlogIn-site articleA Dynamical Model of AI Governability

Early Indicators of Reward Hacking via Reasoning Interpolation

EleutherAI researchers introduce reasoning interpolation, a technique to detect early signs of reward hacking in reinforcement learning. It uses fine-tuned donor models to generate natural exploit-eliciting reasoning prefixes and importance sampling to estimate hack probabilities. While absolute estimates are unreliable early in training, the trend in importance sampling predictions achieves perfect AUC in a controlled setting, suggesting promise as a monitoring signal for RL safety.

EleutherAI BlogIn-site articleEarly Indicators of Reward Hacking via Reasoning Interpolation

Reward Hacking Resarch Update

Interim report on ongoing work on reward hacking

EleutherAI BlogIn-site articleReward Hacking Resarch Update

Pretraining Data Filtering for Open-Weight AI Safety

EleutherAI announces Deep Ignorance, a study showing that filtering pretraining data prevents unsafe knowledge (e.g., biorisk) without losing general performance, and remains robust against fine-tuning attacks. Using string blocklists and ML classifiers, they trained 6.9B models from scratch and found that filtered models resist tampering, though they still learn from in-context information. The paper proposes data filtering as a foundational layer for open-weight model safety.

EleutherAI BlogIn-site articlePretraining Data Filtering for Open-Weight AI Safety

Attention Probes

Attention probes are a novel method for classifying internal states of language models by using an attention layer to aggregate hidden states, avoiding pooling. Multi-head variants (especially 8 heads) outperform mean probes on most datasets, and training code is open-source.

EleutherAI BlogIn-site articleAttention Probes

Research Update: Applications of Local Volume Measurement

EleutherAI researchers tested local volume measurement for detecting model misalignment and anomalous datapoints, found it uncompetitive, and are pivoting to data attribution.

EleutherAI BlogIn-site articleResearch Update: Applications of Local Volume Measurement

Studying inductive biases of random networks via local volumes

In this post, we will study inductive biases of the parameter-function map of random neural networks using star domain volume estimates. This builds on the ideas introduced in Estimating the Probability of Sampling a Trained Neural Network at Random and Neural Redshift: Random Networks are not Random Functions (henceforth NRS).

EleutherAI BlogIn-site articleStudying inductive biases of random networks via local volumes

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

EleutherAI announces the Common Pile v0.1, an 8TB dataset of public domain and openly licensed text, aiming to promote transparency and open science in AI research. The dataset was built in collaboration with multiple institutions, and the trained Comma v0.1 models perform comparably to those trained on unlicensed data.

EleutherAI BlogIn-site articleThe Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

Product Key Memory Sparse Coders

EleutherAI explores using Product Key Memory (PKM) to improve sparse coders. PKM transcoders train faster and are slightly more interpretable than TopK transcoders for moderate expansion factors, though they underperform at extreme scales.

EleutherAI BlogIn-site articleProduct Key Memory Sparse Coders

SAEs trained on the same data don’t learn the same features

Research shows that TopK sparse autoencoders (SAEs) trained on the same data with different random seeds share only about 53% of learned features. Many unshared latents are interpretable. Narrower SAEs have higher feature overlap, while larger SAEs show decreased overlap, consistent with feature splitting and absorption phenomena.

EleutherAI BlogIn-site articleSAEs trained on the same data don’t learn the same features

Partially rewriting an LLM in natural language

This work explores using natural language interpretations of sparse autoencoder (SAE) latents to simulate activations in LLMs. The authors find that current interpretations can identify less than 50% of active latents, and despite high specificity, the extreme imbalance between active and inactive latents leads to many false positives. Predicting activation values from interpretations shows only weak correlations. The results indicate that natural language interpretations are not yet reliable for simulating model activations.

EleutherAI BlogIn-site articlePartially rewriting an LLM in natural language

All sources