Skip to content
AI News HubLIVE
Public articles 11Collected articles 11Trust 82Refresh 60 min
Health Auto-pausedSource type ResearchFull-text rights Full text allowedLast ingested 2026-05-08ID bair-blogStatus Not enabled

Research blog; check individual article license before full text display.

Latest public articles

Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling

Adaptive Parallel Reasoning (APR) empowers LLMs to dynamically decide when to parallelize reasoning, how many threads to spawn, and how to coordinate them. This article analyzes the motivation, methods, training strategies, and open questions in the field.

BAIR BlogIn-site articleAdaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling

Gradient-based Planning for World Models at Longer Horizons

GRASP is a new gradient-based planner for learned dynamics (a world model) that makes long-horizon planning practical by lifting the trajectory into virtual states for parallel optimization, adding stochasticity to state iterates for exploration, and reshaping gradients to avoid brittle state-input gradients through high-dimensional vision models.

BAIR BlogIn-site articleGradient-based Planning for World Models at Longer Horizons

Identifying Interactions at Scale for LLMs

This article presents SPEX and ProxySPEX, algorithms that efficiently identify critical interactions in large language models from three perspectives: feature attribution, data attribution, and mechanistic interpretability. Leveraging structural properties like sparsity, low-degreeness, and hierarchy, these methods discover influential interactions between features, training data, and internal components with fewer ablations, demonstrating strong performance across long contexts, datasets, and model components.

BAIR BlogIn-site articleIdentifying Interactions at Scale for LLMs

Information-Driven Design of Imaging Systems

Researchers have developed a framework to evaluate and optimize imaging systems based on mutual information, predicting performance across four domains and enabling efficient design without task-specific decoders.

BAIR BlogIn-site articleInformation-Driven Design of Imaging Systems

RL without TD learning

This post introduces a reinforcement learning algorithm based on the divide-and-conquer paradigm, which does not rely on temporal difference (TD) learning. The proposed algorithm, Transitive RL (TRL), scales well to long-horizon tasks by recursively splitting trajectories and achieves state-of-the-art performance on challenging OGBench benchmarks without needing to tune the n-step TD hyperparameter.

BAIR BlogIn-site articleRL without TD learning

What exactly does word2vec learn?

A new theory from Berkeley AI Research proves that word2vec reduces to unweighted least-squares matrix factorization, with final representations given by PCA. The model learns discrete orthogonal subspaces sequentially from small initialization, each corresponding to interpretable concepts. The theory predicts features in closed form based on corpus statistics and hyperparameters, matching experiments closely.

BAIR BlogIn-site articleWhat exactly does word2vec learn?

Whole-Body Conditioned Egocentric Video Prediction

BAIR introduces PEVA, a model that predicts egocentric video conditioned on whole-body actions. It uses an autoregressive conditional diffusion transformer trained on Nymeria dataset to simulate atomic actions, long video generation, and visual planning.

BAIR BlogIn-site articleWhole-Body Conditioned Egocentric Video Prediction

Defending against Prompt Injection with Structured Queries (StruQ) and Preference Optimization (SecAlign)

A new BAIR research proposes two fine-tuning defenses against prompt injection attacks, StruQ and SecAlign, which reduce success rates of optimization-free attacks to ~0% and optimization-based attacks to 8%, without additional computational cost or human labor.

BAIR BlogIn-site articleDefending against Prompt Injection with Structured Queries (StruQ) and Preference Optimization (SecAlign)

Repurposing Protein Folding Models for Generation with Latent Diffusion

PLAID is a multimodal generative model that simultaneously generates protein 1D sequence and 3D structure by learning the latent space of protein folding models. It trains on sequence-only data, accepts compositional function and organism prompts, and addresses limitations like all-atom generation, organism specificity, and control specification for practical drug design.

BAIR BlogIn-site articleRepurposing Protein Folding Models for Generation with Latent Diffusion

Scaling Up Reinforcement Learning for Traffic Smoothing: A 100-AV Highway Deployment

We deployed 100 reinforcement learning (RL)-controlled cars into rush-hour highway traffic to smooth congestion and reduce fuel consumption for everyone. Through data-driven simulations, RL agents learned to maximize energy efficiency while maintaining throughput and safety. Field tests show that a small proportion of well-controlled autonomous vehicles (AVs) can significantly improve traffic flow and fuel efficiency, achieving 15-20% energy savings.

BAIR BlogIn-site articleScaling Up Reinforcement Learning for Traffic Smoothing: A 100-AV Highway Deployment

Virtual Personas for Language Models via an Anthology of Backstories

BAIR researchers introduce Anthology, a method for conditioning large language models with detailed personal backstories to create representative, consistent, and diverse virtual personas. This approach outperforms traditional demographic-based conditioning in approximating real human survey responses, offering a cost-effective alternative for social science research.

BAIR BlogIn-site articleVirtual Personas for Language Models via an Anthology of Backstories

All sources