Workflow
Visual Language Models
icon
Search documents
Introducing ExtractBench: The Most Comprehensive Benchmark for Data Extraction from Enterprise Docs
LlamaIndex· 2026-08-11 02:58
Industry Trends and Evaluation Challenges - Enterprise workflows increasingly rely on agents to process unstructured documents, demanding complete, correct, and traceable extraction outputs for meaningful actions[1] - Traditional document extraction benchmarks fail to keep pace with rapid model evolution, typically providing only a single aggregate score over a narrow document set without revealing specific failure causes[1][2][3] Benchmark Performance and System Comparison - ExtractBench evaluates schema extraction in production settings, encompassing **370** enterprise documents across **67** document categories and **8** business domains[2] - Evaluations across **14** systems—including visual language models, open-source pipelines, specialized APIs, and coding agents—reveal that vision model performance collapses on very long documents, and neither vision models nor coding agents provide out-of-the-box visual grounding[4] Company Performance and Technical Advantages - The Llama Extract Agentic+ system achieved the highest overall score of **95.6%**, and is the only system scoring above **90** points on long documents and difficult document scans[5] - The Llama Extract Agentic+ system delivers its performance at a cost that is **one-third** of the next best system[5]