AIAny
Icon for item

Anthropic/BioMysteryBench-full

A collection of biology-focused 'mystery' tasks for benchmarking model performance on biomedical reasoning, evidence synthesis, and problem solving; curated by Anthropic and hosted on Hugging Face, designed for granular evaluation of scientific decision-making.

Introduction

Many benchmarks measure final-answer accuracy but miss whether a model reasons and handles scientific evidence correctly. This dataset assembles multi-step, biology-focused "mystery" tasks to stress-test models' domain reasoning, evidence reconciliation, and practical research-style judgments rather than just surface-level recall.

What Sets It Apart
  • Tasks emphasize multi-step reasoning and evidence handling, so evaluations reveal whether a model reaches conclusions with domain-appropriate justification rather than guessing.
  • Curated by domain-aware designers ( Anthropic ) and provided as a reusable Hugging Face dataset, so it fits into common evaluation pipelines and can be combined with rubrics or automatic judges.
  • Broad artifact support (textual prompts plus supporting artifacts) means prompts often require interpreting auxiliary data, which better approximates real life-science workflows.
Who It's For and Trade-offs
  • Great fit if you evaluate LLMs or multimodal models intended for biomedical literature synthesis, hypothesis ranking, or decision-support in life-science workflows. It helps surface reasoning faults and evidence-misuse risks.
  • Look elsewhere if you only need large-scale classification or simple QA datasets: these tasks are smaller and intentionally harder per example, focusing on quality of reasoning and interpretability over raw throughput or scale.
Where It Fits
  • Complements numeric benchmarks (accuracy-focused) by adding a layer of scientific-validity assessment. Use it alongside automated metrics and human-expert rubrics when assessing model readiness for research-assist roles.

Information

  • Websitehuggingface.co
  • OrganizationsAnthropic, Hugging Face
  • Published date2026/04/29

Categories

More Items

Hugging Face

Provides 315,000 pairwise human-preference votes comparing 15 English TTS models over 300 operational prompts, with 4,500 high‑quality audio renders and structured vote/pair/prompt records for training or evaluating preference/reward models. Metadata under CC-BY-4.0; audio use governed by model providers' terms.

Hugging Face

A 10‑billion‑document retrieval benchmark with per‑document 768‑dim unit‑norm dense embeddings and mGTE sparse embeddings, FineWeb text/metadata, and exact top‑1000 MS MARCO ground truth for ~120k queries. Built for large‑scale evaluation of dense/sparse/hybrid retrieval, filtered search, indexing, ANNS algorithms, and embedding compression.

Hugging Face

Provides 4.5 billion TikTok video records with captions, timestamps, music IDs and engagement counts for research; split across 27 zstd-compressed Parquet files (~289 GB) and sampled via TikTok's mobile API; released for research-use only with privacy and ToS caveats.