AIAny
Icon for item

FineBooks BHL IMPACT Ground Truth

Provides 2,165 historical natural-history page scans paired with ~99.95% expert transcriptions and pixel-aligned PAGE XML layout ground truth for OCR and layout evaluation; multilingual (EN/FR/DE/LA), CC-BY 3.0.

Introduction

Why this matters

High-quality, page-level ground truth for historical print is scarce yet essential for developing and benchmarking OCR and document-layout models. This dataset supplies both near-perfect text transcriptions and precise polygonal layout annotations in the same pixel space as the shipped page images, enabling joint evaluation of text recognition and region detection on real historical material.

What Sets It Apart
  • Dual-purpose ground truth: each page pairs a ~99.95%-accurate reading-order transcription with full PAGE XML polygons, so you can evaluate OCR accuracy and layout/region detection without coordinate transforms.
  • Real historical diversity: 2,165 access pages drawn from six natural-history books (1708–1913) covering English, French, German and Latin (occasional Cyrillic in references), including running text, tables and illustrated plates.
  • Reproducible provenance: the ground truth originates from the IMPACT⇄BHL (2011–2012) collaboration and the dataset bundles the original GT XML, source scandata and the accessioned WebP access images used for mapping.
  • Production-ready packaging: images (WebP) + metadata.parquet + verbatim PAGE XML + derived markdown/docling exports make programmatic workflows (draw boxes, rebuild documents, join catalog metadata) straightforward.
Who it's for and trade-offs

Great fit if you need a moderate-sized, high-fidelity benchmark to train or evaluate OCR/text-recognition and document-layout models on historical print — especially for research on reading-order reconstruction, region detection, or multilingual historical transcription. The CC-BY 3.0 license simplifies reuse with attribution (IMPACT / BHL).

Look elsewhere if you need very large-scale corpora, non-access JP2 masters, or datasets with Gothic/Fraktur type — this corpus contains only antiqua types and includes access-page WebP renditions (≈390 MB). Note also that ~194 pages contain explicit U+FFFD tokens marking illegible glyphs; these are intentional unknown-character markers in the GT, not encoding errors.

Information

  • Websitehuggingface.co
  • Organizationsfinebooks (Hugging Face), IMPACT Centre of Competence, Biodiversity Heritage Library (BHL)
  • Published date2026/06/30

Categories

More Items

Hugging Face

Provides 1.3 billion platform-specific video URLs extracted from CommonCrawl along with crawl metadata (no media included), serving as the source corpus for the LAION-BVD multimodal video dataset; distributed on Hugging Face in Parquet format.

Hugging Face

Provides a public test split of multimodal financial GUI interaction examples for evaluating agents that convert instructions and screenshots into grounded UI actions. Includes step-level screenshots, dialogue history, an OpenAI-style computer_use tool schema, and JSON next-action references; training data available on request.

Hugging Face

Contains 40,000 teacher-generated reasoning traces distilled from the Qwen3.8-27B model for supervised fine-tuning and analysis. Covers code, math, science and logic; each example pairs a <think> chain-of-thought with a final response and is distributed in JSONL/Parquet for SFT workflows.