AIAny
Icon for item

ParseBench

Benchmarks document-parsing systems on real-world enterprise PDFs and images—evaluates tables, charts, content faithfulness, semantic formatting, and visual grounding with human-verified, rule-level tests. Ships with ~2,000 pages, ~169K test rules, and an open evaluation framework for end-to-end pipeline scoring.

Introduction

Why this matters

Enterprise document parsing is a common but brittle dependency for downstream agent workflows: a single mislocated header, dropped sentence, or mislabeled chart value can silently corrupt analytics, auditing, or automated actions. ParseBench targets those precise failure modes by converting human-verified annotations into dense, rule-level tests so teams can see not just overall accuracy but exactly where and how a parser fails in production-like documents.

What Sets It Apart
  • Multi-dimensional diagnostics: instead of a single aggregate score, ParseBench splits evaluation into five capability dimensions (tables, charts, content faithfulness, semantic formatting, visual grounding), each with task-specific metrics that pinpoint structural and semantic faults. So what: you can prioritize fixes (e.g., merged-cell handling vs. reading-order errors) rather than chasing noisy overall gains.
  • Scale and realism: the eval set contains ~2,000 pages from ~1,200 public enterprise documents across insurance, finance, and government, with adversarially hard examples (scans, multi-column layouts, handwritten notes). So what: results better predict real-world failure modes than small synthetic tests.
  • Rule-level granularity and auditing: ~169K human-verified rules (word/sentence/digit omissions, chart datapoint checks, formatting flags, bounding-box grounding) give fine-grained, auditable evidence for regressions. So what: teams can trace a metric drop to specific rule classes and sample pages for targeted remediation.
  • Reproducible evaluation suite: the benchmark includes an evaluation framework and a path to submit leaderboard results (Hugging Face eval-results), enabling comparable, repeatable pipeline scoring across models and toolchains.
Who It's For (Great fit vs. tradeoffs)

Great fit if you: maintain or build document-parsing pipelines for regulated or audit-sensitive domains (finance, insurance, reporting), integrate vision+OCR+NLP components, or need diagnostic signals to prioritize engineering work. ParseBench is designed to reveal edge-case failures that break agentic workflows.

Look elsewhere if you: only need coarse text-extraction quality for casual indexing or search (lighter OCR datasets may be faster), or if your documents are in narrowly constrained, homogeneous formats absent in ParseBench (then a bespoke synthetic test might be more cost-effective).

Where it sits in the toollandscape

ParseBench complements large-scale synthetic OCR/text datasets by focusing on production-like, heterogeneous enterprise documents and on evaluation design (rule-level spot checks and structural metrics) rather than on pretraining data scale. Use it for benchmark-driven improvement and release gating rather than for model pretraining.

How the evaluation works (short)

Each dimension uses specialized metrics: tables use a GTRM (GriTS + TableRecordMatch) structural score; charts verify chart_data_point extraction with configurable tolerances; content faithfulness uses rule-based omission/hallucination checks at word/sentence/digit levels; semantic formatting tests preservation of styles (bold, superscript, strikeout, LaTeX); layout tests grounding via bounding-box, class, and reading-order assertions. Evaluation artifacts are provided as JSONL rule files for reproducible scoring.

Information

  • Websitehuggingface.co
  • Authorsllamaindex (dataset), Boyang Zhang, Sebastián G. Acosta, Preston Carlson, Sacha Bron, Pierre-Loïc Doulcet, Simon Suo
  • Published date2026/04/09

Categories

More Items

Hugging Face

Provides 1.3 billion platform-specific video URLs extracted from CommonCrawl along with crawl metadata (no media included), serving as the source corpus for the LAION-BVD multimodal video dataset; distributed on Hugging Face in Parquet format.

Hugging Face

Provides a public test split of multimodal financial GUI interaction examples for evaluating agents that convert instructions and screenshots into grounded UI actions. Includes step-level screenshots, dialogue history, an OpenAI-style computer_use tool schema, and JSON next-action references; training data available on request.

Hugging Face

Contains 40,000 teacher-generated reasoning traces distilled from the Qwen3.8-27B model for supervised fine-tuning and analysis. Covers code, math, science and logic; each example pairs a <think> chain-of-thought with a final response and is distributed in JSONL/Parquet for SFT workflows.