Why VLMS Just Can
Why forms need a purpose-built parser Representing a form’s structure Finding the boxes Attributing boxes to fields Parsing forms with LlamaParse Try it out A form is one of the most critical types of documents for a bu…
Source profile
AI News Hub tracks LlamaIndex Blog AI updates with visible source status, reuse boundaries, collection method, and published articles.
Official agent and retrieval infrastructure blog; confirm reuse terms before full body display.
Why forms need a purpose-built parser Representing a form’s structure Finding the boxes Attributing boxes to fields Parsing forms with LlamaParse Try it out A form is one of the most critical types of documents for a bu…
What are static embeddings? Attempt #1: MaxSim over raw static tokens Attempt #2: teach the tokens some context It was working? Kinda? Attempt #3: a smarter teacher Attempt #4: skip the teacher, train on the actual metr…
Why we built ExtractBench What ExtractBench evaluates How the ground truth was built Results Reproduce it yourself Enterprise agents increasingly act on critical business data, and much of that data lives in unstructure…
Traditional OCR only transcribes text, while agentic document extraction treats document processing as a reasoning task, using visual grounding, self-correction loops, and plan-act-verify cycles to understand document structure and extract accurate data. This article explains the principles, handling of complex layouts like tables and medical forms, ROI benefits, and implementation best practices.
This article covers the core techniques, workflows, and real-world applications of unstructured data extraction. By leveraging NLP, NER, and LLMs, organizations can automatically extract structured information from vast document repositories, enabling use cases in media monitoring, legal analysis, healthcare research, and more. It also highlights the LlamaParse approach with multi-modal understanding and validation loops, along with best practices and future trends for 2026.
This article breaks down how OCR accuracy is measured (CER, WER, field-level), what factors affect it (resolution, document complexity, handwriting, hardware, condition), how to improve it (pre-processing, synthetic data, LLM post-correction), validation methods, and a 2026 solution landscape. Accuracy is a pipeline problem; the biggest gains come from pre- and post-processing.
AI document classification automates the sorting and tagging of documents, addressing operational bottlenecks at scale. This guide covers how it works, types, use cases, and how to implement a system.
Deep extraction uses an iterative, agent-driven verification loop to achieve near-perfect field accuracy on complex documents, unlike single-pass pipelines that quietly fail on high-volume, high-stakes workflows.
Template OCR works perfectly in demos because the layout never changes, but in real-world scenarios, layout variations cause silent errors. The real alternative is agentic OCR (e.g., LlamaParse) that reads documents by structure, not coordinates, eliminating template maintenance, onboarding delays, and drift. It processes new vendor invoices on the first sight, provides confidence scores for targeted human review, and excels in workflows with high format variability like accounts payable, remittance advice, and logistics.
Mortgage document automation often fails at handoffs between stages, not within stages. Without data provenance, every stage re-verifies numbers, increasing costs and closing times. Agentic extraction, like LlamaParse, provides page-level citations and confidence scoring, enabling targeted human review and tractable audit trails.
LlamaIndex team doubles in size, releases updates: conversational extract, markdown in fast tier, improved tables in Agentic Plus, and 100 users on every plan. AI news includes GPT-5.6 benchmark, Bun rewriting Zig in Rust, Apple suing OpenAI.
LlamaIndex ships multiple product updates in a busy June, including Retrieval Harness for agent file-system tools, LiteParse markdown support, MCP endpoint restructuring, Cost Optimizer improvements, enterprise Usage Tags and User Metadata. Community highlights include a hands-on build with LiteParse and LanceDB, an n8n community node, and several event talks. AI news covers Anthropic's Fable 5 extension, OpenAI DevDay 2026 applications, and more.
LlamaIndex unveils LlamaParse Index with a Retrieval Harness that gives AI agents filesystem-style tools for document traversal, plus visual preservation, managed infra, and observability.
The LlamaParse Platform community node (v5 and v6) is now an officially verified n8n community node. It exposes five LlamaCloud resources (Parse, Classify, Split, Extract, Retrieve) that can be used as tools in n8n AI Agents. v5 rewrote the foundation with direct HTTP calls and configurable API base URL. v6 consolidated multiple nodes into one and added index actions. The post presents three example workflows: retrievers as agent tools, a classify-extract-verify pipeline, and evaluating parsed outputs across different parsing modes.
LiteParse v2.1 introduces the fastest open-source, model-free PDF-to-markdown pipeline, achieving top scores on three benchmarks and offering speed and portability across multiple runtimes.
This article details how the authors improved their LiteParse document parsing skill for Claude agents through iterative evaluation, trace analysis, and optimization, achieving a 37% cost reduction and higher answer quality. Key anti-patterns like re-parsing, unnecessary OCR, and excessive grep turns were identified and fixed.
Major updates include ParseBench at CVPR 2026, Parse-Flow for visual document intelligence, Anthropic Fable 5 benchmark results, new Granular Bounding Boxes in LlamaParse, and The Agent Open pickleball tournament.
This article explores the true meaning of PDF searchability. Quick OCR methods like Adobe Acrobat and free online tools work for clean documents but fail on tables, multi-column layouts, and poor scans. Even a 95% accurate text layer leaves errors that cause searches to miss targets. For large-scale or AI-driven processing, structured output from tools like LlamaParse is necessary to preserve reading order and table structure. True searchability depends on accuracy and structure, not just the presence of a text layer.
Organizations face significant challenges in extracting structured metadata from complex legal contracts due to variability in language, structure, and formatting. Modern systems combine layout-aware parsing, machine learning, semantic extraction, and schema mapping to transform unstructured legal agreements into machine-readable data. LlamaParse offers a structured platform integrating these capabilities for production workflows.
Parse-Flow is an open-source project that combines a visual workflow designer, async worker, and live event dashboard to orchestrate document processing primitives — parsing, extraction, classification, and splitting. Built on llama-agents workflows, it uses Redis and Postgres for job queuing and event persistence, with a three-step state machine (bootstrap, worker, router) that interprets user-defined flows defined in JSON. This article details its architecture, design rationale, and the importance of robust document intelligence for enterprise AI systems.
This post compares grep (lexical search) and RAG (semantic search) for AI agents. Grep is fast and precise on small plain-text corpora but cannot handle unstructured documents and doesn't scale. RAG solves scalability via parsing, chunking, embedding, and vector indexing, enabling vocabulary-agnostic search. The recommended approach is layered: parse unstructured documents, use semantic search at scale, and keep grep for suitable cases.
This week's LlamaIndex newsletter highlights ParseBench, the first OCR benchmark for AI agents, along with new open-source tools: a sandboxed CLI agent for secure document interactions and a self-hostable document parsing server. Community events in Singapore and NYC are also recapped.
This post walks through a demo AI agent that ingests SEC filings, searches across them, and answers questions with precise citations that highlight the exact source text on the original PDF page. The key ingredient is LiteParse, which extracts text along with bounding box coordinates. The project uses a simple keyword search instead of vector databases, and integrates SEC EDGAR for direct filing retrieval.
Mortgage document automation leverages intelligent document processing to transform document-heavy workflows into structured, machine-driven processes, improving efficiency and reducing errors. This article examines the complexities of mortgage document handling, the automation workflow (ingestion, classification, extraction, validation, human review, and system integration), challenges, and best practices for implementation with LlamaParse.
This article examines the shortcomings of standard OCR in KYC workflows, including its inability to handle real-world identity documents with security features, varied layouts, and non-Latin scripts. It introduces agentic OCR (e.g., LlamaParse) that uses layout-aware segmentation, model orchestration, and self-correction loops to achieve over 90-95% straight-through processing, and discusses implications across banking, insurance, and crypto industries.
This week's LlamaIndex newsletter highlights intelligent table extraction, LiteSearch for local document retrieval, improved Word document processing, and integrations with Gemini Live API, along with guides for legal discovery and community projects.
This edition introduces ParseBench, the first OCR benchmark designed for AI agents, along with LiteParse's explosive growth, structure-aware PDF QA pipeline, VLM-powered OCR production insights, NYC fintech workshop, and secure document agents.
Highlights include the launch of ParseBench, the first document OCR benchmark for AI agents; LiteParse officially joining the LlamaIndex ecosystem; comprehensive benchmarking of Anthropic's Opus 4.7; and an upcoming NYC FinTech Week AI event.
The refactored LlamaParse Platform MCP shifts focus to document processing with Parse, Classify, and Split services. This post covers the MCP tools, connection methods, OAuth authentication, file upload solutions, observability, rate limiting, and deployment considerations.
liteparse-server is a self-hosted HTTP API wrapping the LiteParse document parsing engine, supporting PDFs, Office documents, and images with precise spatial layout text extraction and OCR. It addresses the latency, cost, and privacy concerns of cloud parsing, suitable for RAG and vision model workflows. Two deployment modes: slim (no dependencies) and full stack (with Redis caching, rate limiting, OpenTelemetry tracing, and Prometheus metrics).