Use Case Use Case: Manufacturing Industry Sep 1, 2026 Use Case Use Case: Manufacturing Industry Sep 1, 2026 Authors Unstructured In this article Join our newsletter to receive updates about our features. In this article…
Use Case Use Case: Insurance Industry Sep 1, 2026 Use Case Use Case: Insurance Industry Sep 1, 2026 Authors Unstructured In this article Join our newsletter to receive updates about our features. In this article Structu…
Unstructured introduces Transform MCP, a tool that lets AI agents process documents as a tool call. It handles 60+ file types, provides structured output, and integrates seamlessly with MCP clients. Key benefits include composability, inspectability, reproducibility, and portability. Available now with a free tier of 15,000 pages per month.
This article introduces how Unstructured Platform handles over 60 unstructured data formats using a multi-layered parsing strategy, including rule-based parsers, OCR, and VLMs. It also discusses the trade-off in MCP server tool count and how to optimize context window usage.
Object Detection (OD) remains a critical foundation for document transformation pipelines, even with the rise of VLMs. This article explains why fine-tuning OD models is far from plug-and-play, covering challenges such as training framework bugs, inconsistent label philosophies across datasets, and catastrophic forgetting. Unstructured's approach, building on IBM Heron, achieved significant accuracy gains through careful fine-tuning.
The article discusses how AI agents fail because they cannot access the majority of unstructured data within organizations. It introduces Unstructured as a platform that processes over 65 file types, extracts, chunks, enriches, and embeds data into Databricks lakehouses, enabling agents to access full business context while maintaining governance through Unity Catalog.
The Naval Sea Systems Command awarded Unstructured a contract to design an AI-enabled solution that helps warfighters surface mission-critical information faster, reduce operator workload, and accelerate decision-making in Anti-Submarine Warfare and surface warfare operations. The solution integrates Unstructured's data ingestion with Elastic's enterprise search to mine heterogeneous data sources, initially deployed on CV-TSC and USW-DSS systems with future applicability to JADC2 and C5ISR.
Unstructured introduces Extract, a new enrichment node that extracts structured JSON data from documents using LLM or regex, enabling intelligent document processing within existing workflows.
Unstructured launches webhooks to automate downstream actions based on job lifecycle events, allowing integration with any endpoint via workspace or workflow scopes.
Unstructured found that combining high-quality datasets with incompatible annotation styles degraded model performance. They built an agentic harmonization pipeline using a VLM to reconcile label differences, resulting in improved metrics across 14 of 17 benchmarks.
Unstructured's SCORE-Bench benchmark evaluates five frontier models on enterprise document parsing, revealing a significant gap between raw model calls and optimized pipelines. While models excel in reasoning and hallucination control (especially Claude Opus 4.6), they lag by up to 23 percentage points in table extraction, document structure, and output consistency. The gap is attributed to configuration rather than capability, and can be closed via optimized prompting, post-processing, and output structure enforcement.
Unstructured's new guide on advanced Retrieval-Augmented Generation (RAG) techniques, covering smart chunking, metadata filtering, GraphRAG, hybrid search, and agentic workflows, aimed at building scalable enterprise AI pipelines.
Unstructured announces major updates including a simplified drag-and-drop interface, Generative Refinement for higher fidelity outputs, and simplified pricing with a free tier. The new workflow combines high-resolution partitioning with VLM-powered enrichments to achieve superior accuracy and structure preservation.