Skip to content
AI News HubLIVE
Public articles 32Collected articles 38Trust 84Refresh 120 min
Health HealthySource type OfficialFull-text rights Official full textLast ingested 2026-09-28ID llamaindex-blogStatus Enabled

Official agent and retrieval infrastructure blog; confirm reuse terms before full body display.

Latest public articles

Why VLMS Just Can

Why forms need a purpose-built parser Representing a form’s structure Finding the boxes Attributing boxes to fields Parsing forms with LlamaParse Try it out A form is one of the most critical types of documents for a bu…

LlamaIndex BlogIn-site articleWhy VLMS Just Can

Exploring Static Embedding Retrieval

What are static embeddings? Attempt #1: MaxSim over raw static tokens Attempt #2: teach the tokens some context It was working? Kinda? Attempt #3: a smarter teacher Attempt #4: skip the teacher, train on the actual metr…

LlamaIndex BlogIn-site articleExploring Static Embedding Retrieval

How Agentic Document Extraction Improves Accuracy and Automation

Traditional OCR only transcribes text, while agentic document extraction treats document processing as a reasoning task, using visual grounding, self-correction loops, and plan-act-verify cycles to understand document structure and extract accurate data. This article explains the principles, handling of complex layouts like tables and medical forms, ROI benefits, and implementation best practices.

LlamaIndex BlogIn-site articleHow Agentic Document Extraction Improves Accuracy and Automation

Unstructured Data Extraction: Turn Documents into Insights

This article covers the core techniques, workflows, and real-world applications of unstructured data extraction. By leveraging NLP, NER, and LLMs, organizations can automatically extract structured information from vast document repositories, enabling use cases in media monitoring, legal analysis, healthcare research, and more. It also highlights the LlamaParse approach with multi-modal understanding and validation loops, along with best practices and future trends for 2026.

LlamaIndex BlogIn-site articleUnstructured Data Extraction: Turn Documents into Insights

OCR Accuracy Explained: How to Improve It

This article breaks down how OCR accuracy is measured (CER, WER, field-level), what factors affect it (resolution, document complexity, handwriting, hardware, condition), how to improve it (pre-processing, synthetic data, LLM post-correction), validation methods, and a 2026 solution landscape. Accuracy is a pipeline problem; the biggest gains come from pre- and post-processing.

LlamaIndex BlogIn-site articleOCR Accuracy Explained: How to Improve It

AI Document Classification: A Practical Guide

AI document classification automates the sorting and tagging of documents, addressing operational bottlenecks at scale. This guide covers how it works, types, use cases, and how to implement a system.

LlamaIndex BlogIn-site articleAI Document Classification: A Practical Guide

Why Deep Extraction is Superior to Single-Pass Pipelines

Deep extraction uses an iterative, agent-driven verification loop to achieve near-perfect field accuracy on complex documents, unlike single-pass pipelines that quietly fail on high-volume, high-stakes workflows.

LlamaIndex BlogIn-site articleWhy Deep Extraction is Superior to Single-Pass Pipelines

The Real Alternative to Template OCR Isn

Template OCR works perfectly in demos because the layout never changes, but in real-world scenarios, layout variations cause silent errors. The real alternative is agentic OCR (e.g., LlamaParse) that reads documents by structure, not coordinates, eliminating template maintenance, onboarding delays, and drift. It processes new vendor invoices on the first sight, provides confidence scores for targeted human review, and excels in workflows with high format variability like accounts payable, remittance advice, and logistics.

LlamaIndex BlogIn-site articleThe Real Alternative to Template OCR Isn

Mortgage Banking Document Automation: Common Workflow Gaps

Mortgage document automation often fails at handoffs between stages, not within stages. Without data provenance, every stage re-verifies numbers, increasing costs and closing times. Agentic extraction, like LlamaParse, provides page-level citations and confidence scoring, enabling targeted human review and tractable audit trails.

LlamaIndex BlogIn-site articleMortgage Banking Document Automation: Common Workflow Gaps

Parse, Extract, Classify — now each with its own MCP | Weekly Newsletter

LlamaIndex team doubles in size, releases updates: conversational extract, markdown in fast tier, improved tables in Agentic Plus, and 100 users on every plan. AI news includes GPT-5.6 benchmark, Bun rewriting Zig in Rust, Apple suing OpenAI.

LlamaIndex BlogIn-site articleParse, Extract, Classify — now each with its own MCP | Weekly Newsletter

Parse, Extract, Classify — now each with its own MCP | Weekly Newsletter

LlamaIndex ships multiple product updates in a busy June, including Retrieval Harness for agent file-system tools, LiteParse markdown support, MCP endpoint restructuring, Cost Optimizer improvements, enterprise Usage Tags and User Metadata. Community highlights include a hands-on build with LiteParse and LanceDB, an n8n community node, and several event talks. AI news covers Anthropic's Fable 5 extension, OpenAI DevDay 2026 applications, and more.

LlamaIndex BlogIn-site articleParse, Extract, Classify — now each with its own MCP | Weekly Newsletter

LlamaParse Retrieval Harness: Filesystem Primitives for AI Agents

LlamaIndex unveils LlamaParse Index with a Retrieval Harness that gives AI agents filesystem-style tools for document traversal, plus visual preservation, managed infra, and observability.

LlamaIndex BlogIn-site articleLlamaParse Retrieval Harness: Filesystem Primitives for AI Agents

LlamaParse Platform Node for n8n: Parse, Classify, Extract & Retrieve Documents with AI

The LlamaParse Platform community node (v5 and v6) is now an officially verified n8n community node. It exposes five LlamaCloud resources (Parse, Classify, Split, Extract, Retrieve) that can be used as tools in n8n AI Agents. v5 rewrote the foundation with direct HTTP calls and configurable API base URL. v6 consolidated multiple nodes into one and added index actions. The post presents three example workflows: retrievers as agent tools, a classify-extract-verify pipeline, and evaluating parsed outputs across different parsing modes.

LlamaIndex BlogIn-site articleLlamaParse Platform Node for n8n: Parse, Classify, Extract & Retrieve Documents with AI

Markdown Comes to LiteParse

LiteParse v2.1 introduces the fastest open-source, model-free PDF-to-markdown pipeline, achieving top scores on three benchmarks and offering speed and portability across multiple runtimes.

LlamaIndex BlogIn-site articleMarkdown Comes to LiteParse

Building a Faster, Cheaper PDF-Parsing Skill for Claude Agents: A LiteParse Case Study

This article details how the authors improved their LiteParse document parsing skill for Claude agents through iterative evaluation, trace analysis, and optimization, achieving a 37% cost reduction and higher answer quality. Key anti-patterns like re-parsing, unnecessary OCR, and excessive grep turns were identified and fixed.

LlamaIndex BlogIn-site articleBuilding a Faster, Cheaper PDF-Parsing Skill for Claude Agents: A LiteParse Case Study

LlamaIndex Newsletter 6-10-26

Major updates include ParseBench at CVPR 2026, Parse-Flow for visual document intelligence, Anthropic Fable 5 benchmark results, new Granular Bounding Boxes in LlamaParse, and The Agent Open pickleball tournament.

LlamaIndex BlogIn-site articleLlamaIndex Newsletter 6-10-26

How to Make a PDF Searchable: Methods and Limits

This article explores the true meaning of PDF searchability. Quick OCR methods like Adobe Acrobat and free online tools work for clean documents but fail on tables, multi-column layouts, and poor scans. Even a 95% accurate text layer leaves errors that cause searches to miss targets. For large-scale or AI-driven processing, structured output from tools like LlamaParse is necessary to preserve reading order and table structure. True searchability depends on accuracy and structure, not just the presence of a text layer.

LlamaIndex BlogIn-site articleHow to Make a PDF Searchable: Methods and Limits

Extract Contract Metadata: Methods, Challenges, and Workflows

Organizations face significant challenges in extracting structured metadata from complex legal contracts due to variability in language, structure, and formatting. Modern systems combine layout-aware parsing, machine learning, semantic extraction, and schema mapping to transform unstructured legal agreements into machine-readable data. LlamaParse offers a structured platform integrating these capabilities for production workflows.

LlamaIndex BlogIn-site articleExtract Contract Metadata: Methods, Challenges, and Workflows

Parse-Flow: Open-Source Visual Document Intelligence Workflow Designer

Parse-Flow is an open-source project that combines a visual workflow designer, async worker, and live event dashboard to orchestrate document processing primitives — parsing, extraction, classification, and splitting. Built on llama-agents workflows, it uses Redis and Postgres for job queuing and event persistence, with a three-step state machine (bootstrap, worker, router) that interprets user-defined flows defined in JSON. This article details its architecture, design rationale, and the importance of robust document intelligence for enterprise AI systems.

LlamaIndex BlogIn-site articleParse-Flow: Open-Source Visual Document Intelligence Workflow Designer

grep vs. RAG: Choosing the Right Search Strategy for AI Agents

This post compares grep (lexical search) and RAG (semantic search) for AI agents. Grep is fast and precise on small plain-text corpora but cannot handle unstructured documents and doesn't scale. RAG solves scalability via parsing, chunking, embedding, and vector indexing, enabling vocabulary-agnostic search. The recommended approach is layered: parse unstructured documents, use semantic search at scale, and keep grep for suitable cases.

LlamaIndex BlogIn-site articlegrep vs. RAG: Choosing the Right Search Strategy for AI Agents

LlamaIndex Newsletter 5-19-26

This week's LlamaIndex newsletter highlights ParseBench, the first OCR benchmark for AI agents, along with new open-source tools: a sandboxed CLI agent for secure document interactions and a self-hostable document parsing server. Community events in Singapore and NYC are also recapped.

LlamaIndex BlogIn-site articleLlamaIndex Newsletter 5-19-26

How to Build a Financial Due Diligence Agent with LiteParse

This post walks through a demo AI agent that ingests SEC filings, searches across them, and answers questions with precise citations that highlight the exact source text on the original PDF page. The key ingredient is LiteParse, which extracts text along with bounding box coordinates. The project uses a simple keyword search instead of vector databases, and integrates SEC EDGAR for direct filing retrieval.

LlamaIndex BlogIn-site articleHow to Build a Financial Due Diligence Agent with LiteParse

Mortgage Document Automation: Transforming Loan Processing

Mortgage document automation leverages intelligent document processing to transform document-heavy workflows into structured, machine-driven processes, improving efficiency and reducing errors. This article examines the complexities of mortgage document handling, the automation workflow (ingestion, classification, extraction, validation, human review, and system integration), challenges, and best practices for implementation with LlamaParse.

LlamaIndex BlogIn-site articleMortgage Document Automation: Transforming Loan Processing

OCR for KYC: Why Standard Text Extraction Falls Short

This article examines the shortcomings of standard OCR in KYC workflows, including its inability to handle real-world identity documents with security features, varied layouts, and non-Latin scripts. It introduces agentic OCR (e.g., LlamaParse) that uses layout-aware segmentation, model orchestration, and self-correction loops to achieve over 90-95% straight-through processing, and discusses implications across banking, insurance, and crypto industries.

LlamaIndex BlogIn-site articleOCR for KYC: Why Standard Text Extraction Falls Short

LlamaIndex Newsletter: Intelligent Table Extraction & LiteSearch

This week's LlamaIndex newsletter highlights intelligent table extraction, LiteSearch for local document retrieval, improved Word document processing, and integrations with Gemini Live API, along with guides for legal discovery and community projects.

LlamaIndex BlogIn-site articleLlamaIndex Newsletter: Intelligent Table Extraction & LiteSearch

LlamaIndex Newsletter 2026-04-14

This edition introduces ParseBench, the first OCR benchmark designed for AI agents, along with LiteParse's explosive growth, structure-aware PDF QA pipeline, VLM-powered OCR production insights, NYC fintech workshop, and secure document agents.

LlamaIndex BlogIn-site articleLlamaIndex Newsletter 2026-04-14

LlamaIndex Newsletter 2026-04-21

Highlights include the launch of ParseBench, the first document OCR benchmark for AI agents; LiteParse officially joining the LlamaIndex ecosystem; comprehensive benchmarking of Anthropic's Opus 4.7; and an upcoming NYC FinTech Week AI event.

LlamaIndex BlogIn-site articleLlamaIndex Newsletter 2026-04-21

LlamaParse MCP: Agentic OCR tools for your AI agents

The refactored LlamaParse Platform MCP shifts focus to document processing with Parse, Classify, and Split services. This post covers the MCP tools, connection methods, OAuth authentication, file upload solutions, observability, rate limiting, and deployment considerations.

LlamaIndex BlogIn-site articleLlamaParse MCP: Agentic OCR tools for your AI agents

Introducing liteparse-server: Self-Hosted Document Parsing and OCR for AI Workflows

liteparse-server is a self-hosted HTTP API wrapping the LiteParse document parsing engine, supporting PDFs, Office documents, and images with precise spatial layout text extraction and OCR. It addresses the latency, cost, and privacy concerns of cloud parsing, suitable for RAG and vision model workflows. Two deployment modes: slim (no dependencies) and full stack (with Redis caching, rate limiting, OpenTelemetry tracing, and Prometheus metrics).

LlamaIndex BlogIn-site articleIntroducing liteparse-server: Self-Hosted Document Parsing and OCR for AI Workflows

All sources