Skip to content
AI News HubLIVE
Public articles 11Collected articles 12Trust 88Refresh 720 min
Health HealthySource type ResearchFull-text rights Full text allowedLast ingested 2026-07-07ID lilian-wengStatus Enabled

Public independent AI research blog; verify individual post license before full body display.

Latest public articles

Harness Engineering for Self-Improvement

The concept of recursive self-improvement (RSI) dates back to I. J. Good (1965), where he defined an “ultraintelligent machine” as a system that can surpass humans in all intellectual activities and design better machines to improve itself. Yudkowsky (2008) used the phrase “recursive self-improvement” for a specific feedback loop: an AI uses its current intelligence to improve the cognitive machinery that produces its intelligence. This feedback loop in modern AI may indicate the model rewriting its own weights directly, or more broadly the model improves the training pipeline and the deployment system, which in turn enables a better successor model with improved performance across economically valuable tasks. The speed of research development in AI has been shown to drastically accelerat…

Lilian WengIn-site articleHarness Engineering for Self-Improvement

Scaling Laws, Carefully

Scaling laws are one of the most critical empirical findings in deep learning, describing power-law relationships between model size, data, compute, and loss. This article reviews the development from early theory to modern empirical studies, including Kaplan et al.'s classic scaling laws and the Chinchilla scaling laws, and discusses key findings such as compute-optimal allocation.

Lilian WengIn-site articleScaling Laws, Carefully

Reward Hacking in Reinforcement Learning

Reward hacking occurs when a reinforcement learning agent exploits flaws or ambiguities in the reward function to achieve high rewards without genuinely learning or completing the intended task. With the rise of language models and RLHF, reward hacking has become a critical practical challenge. This article covers the definition, types, causes, and potential mitigations of reward hacking.

Lilian WengIn-site articleReward Hacking in Reinforcement Learning

Extrinsic Hallucinations in LLMs

This article by Lilian Weng focuses on extrinsic hallucinations in large language models, where models generate fabricated content not grounded in provided context or world knowledge. It explores causes such as pre-training data issues and fine-tuning new knowledge, discusses detection methods including retrieval-augmented evaluation and sampling-based approaches, and presents anti-hallucination techniques like RAG, chain-of-verification, sampling adjustments, and fine-tuning for factuality and attribution.

Lilian WengIn-site articleExtrinsic Hallucinations in LLMs

Diffusion Models for Video Generation

Diffusion models have demonstrated strong results on image synthesis in past years. Now the research community has started working on a harder task—using it for video generation. The task itself is a superset of the image case, since an image is a video of 1 frame, and it is much more challenging because: it has extra requirements on temporal consistency across frames in time, which naturally demands more world knowledge to be encoded into the model; and in comparison to text or images, it is more difficult to collect large amounts of high-quality, high-dimensional video data, let alone text-video pairs.

Lilian WengIn-site articleDiffusion Models for Video Generation

Thinking about High-Quality Human Data

High-quality data is the fuel for modern deep learning model training. This article explores how to collect high-quality data through human annotation, including task design, rater selection and training, data aggregation, and quality assurance. It covers the wisdom of the crowd, methods for measuring rater agreement (e.g., Cohen's Kappa, MACE), and two annotation paradigms (descriptive vs. prescriptive). Additionally, it discusses techniques to identify mislabeled data using influence functions, training dynamics (e.g., data maps, forgetting events, AUM), and noisy cross-validation.

Lilian WengIn-site articleThinking about High-Quality Human Data

Adversarial Attacks on LLMs

A comprehensive survey of adversarial attacks on large language models, covering threat models, attack types including token manipulation, gradient-based attacks, jailbreak prompting, and red-teaming techniques. The article discusses the challenges and methods for both black-box and white-box settings.

Lilian WengIn-site articleAdversarial Attacks on LLMs

LLM Powered Autonomous Agents

This article explores autonomous agents powered by large language models (LLMs) as their core controller. The system comprises three main components: planning (task decomposition and self-reflection), memory (short-term via in-context learning, long-term via external vector stores), and tool use (calling external APIs). It covers case studies like ChemCrow and Generative Agents, proof-of-concepts such as AutoGPT, GPT-Engineer, and BabyAGI, and discusses challenges like finite context windows.

Lilian WengIn-site articleLLM Powered Autonomous Agents

Prompt Engineering

This article provides a comprehensive overview of prompt engineering methods for large language models, covering basic prompting, instruction prompting, self-consistency sampling, chain-of-thought prompting, automatic prompt design, and augmented language models.

Lilian WengIn-site articlePrompt Engineering

The Transformer Family Version 2.0

This article is a major update to Lilian Weng's 2020 post on the Transformer family, doubling its length. It systematically reviews numerous recent improvements to the Transformer architecture, covering attention mechanisms, positional encoding, long-context support, adaptive modeling, and efficient attention, including the latest advances such as Transformer-XL, Rotary position embedding, ALiBi, and the Universal Transformer.

Lilian WengIn-site articleThe Transformer Family Version 2.0

Large Transformer Model Inference Optimization

A comprehensive overview of techniques to optimize inference for large transformer models, including distillation, quantization, pruning, sparsity, mixture-of-experts, and architectural improvements. The article discusses challenges such as memory footprint and low parallelizability, and presents methods to reduce memory usage, computation, and latency.

Lilian WengIn-site articleLarge Transformer Model Inference Optimization

All sources