AI expert Chip Huyen outlines six common pitfalls when building generative AI applications: using GenAI unnecessarily, confusing bad product with bad AI, starting too complex, over-indexing on early success, forgoing human evaluation, and crowdsourcing use cases without strategy. The article provides real-world examples and practical advice to avoid these mistakes.
Intelligent agents are considered the ultimate goal of AI. This article from AI Engineering (2025) provides a comprehensive framework for understanding agents, focusing on tools and planning. It covers agent overview, tool categories (knowledge augmentation, capability extension, write actions), planning processes (generation, validation, execution, reflection), and evaluation of agent failures.
After studying how companies deploy generative AI applications, this post outlines the common components of a generative AI platform. Starting from a simple query-response architecture, it progressively adds context enhancement (RAG, query rewriting), guardrails (input/output), model routers and gateways, caching (prompt, exact, semantic), complex logic and write actions, and observability/orchestration. Each component's trade-offs and implementation considerations are discussed.
The author explores three metrics for personal growth: rate of change, time to solve problems, and number of future options. Inspired by discussions with friends, she proposes heuristics that prioritize novelty and exploration over traditional measures like net worth.
Chip Huyen analyzes nearly 900 popular open-source AI projects, revealing explosive growth in applications and AI engineering layers in 2023, while infrastructure remained relatively stable. The Chinese open-source ecosystem is diverging significantly from the West, with many Chinese-focused models and tools emerging.
This article explores predicting user preferences for AI model responses to enable model routing and improve efficiency. The author demonstrates that preference prediction is feasible with a small amount of data and shows its performance across different prompts.
An in-depth exploration of sampling strategies including temperature, top-k, top-p, test time compute, and structured outputs, explaining how they affect the creativity, consistency, and reliability of AI model responses.
This comprehensive article explores multimodal AI systems, particularly Large Multimodal Models (LMMs). It covers the rationale for multimodality, data modalities, multimodal tasks, and dives into the architectures and training of CLIP and Flamingo models. The post also discusses active research directions such as generating multimodal outputs, instruction-following, and efficient adapters.
This article summarizes ten major research directions in large language models, covering hallucinations, context learning, multimodality, speed and cost, new architectures, GPU alternatives, agents, human preference learning, chat interface efficiency, and non-English language models. Based on discussions with industry and academia, the author analyzes the current state and challenges of each direction.
Chip Huyen's talk at Fully Connected provides a simple framework for teams to navigate generative AI strategy, born from conversations with friends struggling to define their approach. She plans to expand the talk into a full article while welcoming community feedback.