What We Learned Trying to Catch AI Liars: An Aletheia's Quest Retrospective
What we learned while building black-box and white-box detectors for AI deception during Aletheia's Quest.
Source profile
AI News Hub tracks EleutherAI Blog AI updates with visible source status, reuse boundaries, collection method, and published articles.
Official research collective blog; confirm reuse terms before full body display.
What we learned while building black-box and white-box detectors for AI deception during Aletheia's Quest.
A toy dynamical model of whether the AI workforce that builds future AI ends up cooperative or uncooperative: where the basin boundary lies, what current evidence says about which side we are on, and what would tell us we are on the good path.
EleutherAI researchers introduce reasoning interpolation, a technique to detect early signs of reward hacking in reinforcement learning. It uses fine-tuned donor models to generate natural exploit-eliciting reasoning prefixes and importance sampling to estimate hack probabilities. While absolute estimates are unreliable early in training, the trend in importance sampling predictions achieves perfect AUC in a controlled setting, suggesting promise as a monitoring signal for RL safety.
Interim report on ongoing work on reward hacking
EleutherAI announces Deep Ignorance, a study showing that filtering pretraining data prevents unsafe knowledge (e.g., biorisk) without losing general performance, and remains robust against fine-tuning attacks. Using string blocklists and ML classifiers, they trained 6.9B models from scratch and found that filtered models resist tampering, though they still learn from in-context information. The paper proposes data filtering as a foundational layer for open-weight model safety.
Attention probes are a novel method for classifying internal states of language models by using an attention layer to aggregate hidden states, avoiding pooling. Multi-head variants (especially 8 heads) outperform mean probes on most datasets, and training code is open-source.
EleutherAI researchers tested local volume measurement for detecting model misalignment and anomalous datapoints, found it uncompetitive, and are pivoting to data attribution.
In this post, we will study inductive biases of the parameter-function map of random neural networks using star domain volume estimates. This builds on the ideas introduced in Estimating the Probability of Sampling a Trained Neural Network at Random and Neural Redshift: Random Networks are not Random Functions (henceforth NRS).
EleutherAI announces the Common Pile v0.1, an 8TB dataset of public domain and openly licensed text, aiming to promote transparency and open science in AI research. The dataset was built in collaboration with multiple institutions, and the trained Comma v0.1 models perform comparably to those trained on unlicensed data.
EleutherAI explores using Product Key Memory (PKM) to improve sparse coders. PKM transcoders train faster and are slightly more interpretable than TopK transcoders for moderate expansion factors, though they underperform at extreme scales.
Research shows that TopK sparse autoencoders (SAEs) trained on the same data with different random seeds share only about 53% of learned features. Many unshared latents are interpretable. Narrower SAEs have higher feature overlap, while larger SAEs show decreased overlap, consistent with feature splitting and absorption phenomena.
This work explores using natural language interpretations of sparse autoencoder (SAE) latents to simulate activations in LLMs. The authors find that current interpretations can identify less than 50% of active latents, and despite high specificity, the extreme imbalance between active and inactive latents leads to many false positives. Predicting activation values from interpretations shows only weak correlations. The results indicate that natural language interpretations are not yet reliable for simulating model activations.