Introducing ExtractBench: The Most Comprehensive Benchmark for Data Extraction from Enterprise Docs
LlamaIndex·2026-08-10 19:55

Industry Trends & Benchmark Limitations - Enterprise workflows increasingly rely on agents to act on unstructured documents, requiring extraction outputs to be complete, correct, and traceable [1] - Existing benchmarks often provide only a single aggregate score over a narrow set of documents, failing to diagnose specific failure causes such as dropped rows or mislabeled fields [2][3] Evaluation Scope & System Performance - Extract Bench evaluates schema-guided extraction across 370 enterprise documents spanning 67 document categories and 8 business domains [2] - Evaluations across 14 systems reveal that vision models experience performance collapses on very long documents and lack out-of-the-box grounding [4] - Neither vision models nor coding agents return grounding natively without human intervention [4] Product Performance & Cost Efficiency - Llama Extract Agentic Plus achieves the highest overall score at 95.6% [5] - Llama Extract Agentic Plus is the sole system scoring above 90 points on long documents and hard document scans [5] - Llama Extract Agentic Plus operates at a cost equivalent to one-third of the next best system [5]

Introducing ExtractBench: The Most Comprehensive Benchmark for Data Extraction from Enterprise Docs - Reportify