AIAny
Icon for item

Xperience-10M

Provides 10 million synchronized egocentric experience episodes with structured 3D/4D multimodal annotations — 2.88B RGB frames, 720M depth frames, 576M pose/mocap frames and ~1PB total. Designed for embodied AI, robotics, and multimodal pretraining; research-only, gated access.

Introduction

Most embodied-AI and robot-learning gaps come from a lack of large-scale, real-world streams that jointly capture vision, motion, audio and spatial geometry. This dataset converts everyday first‑person interactions into a unified 4D training corpus so models can learn perception, dynamics, and interaction from the same synchronized experience stream.

What Sets It Apart
  • Truly multimodal and large-scale: 10M interaction episodes with synchronized vision (four fisheye streams), audio, stereo/monocular depth, IMU, hand and full-body mocap, and camera poses — totals on the order of 2.88B RGB frames, 720M depth frames, and ~1PB of data. This scale supports pretraining across temporal, spatial and kinematic modalities.
  • Structured 3D/4D annotations: per-episode annotation files include calibration, geometry, trajectories, hierarchical natural-language captions, object instances and dense pose/mocap — enabling cross-modal grounding (language↔action↔3D) and long-horizon trajectory learning.
  • Egocentric, in-the-wild focus: first-person captures of human interactions emphasize human-object interaction, manipulation, and embodied behavior rather than lab-constrained, staged scenes — useful for imitation learning, real-to-sim transfer, and world models that must reason about agent-centric observations.
  • Designed for downstream embodied tasks: the data supports SLAM/pose estimation, action recognition/localization, multimodal pretraining (vision+language+motion), sim-to-real pipelines, and robotics imitation learning with dense kinematic supervision.
Who it's for — and tradeoffs

Great fit if you need large-scale, synchronized multimodal experience traces for training embodied or spatially-aware models (e.g., multimodal foundation models, robot policies, or 3D reconstruction systems). The dataset's scale and annotation density make it especially valuable for pretraining or for supervised tasks requiring motion/pose labels aligned with video and language.

Look elsewhere if you need fully open commercial use or very low-bandwidth samples: access is research-only and gated, the dataset is enormous (~1PB) so storage and compute costs are substantial, and many users will rely on provided samples rather than the full corpus. Also, because captures are egocentric, tasks requiring third‑person multi-view cinematic footage are not the primary fit.

Practical notes
  • Language annotations are in English and vocabulary/annotation schemas are hierarchical for multi-granularity supervision.
  • The dataset is released for non-commercial research use and may require an access agreement.
  • Typical uses: embodied model pretraining, action-language grounding, 3D/4D reconstruction and tracking, imitation learning and sensor-fusion research.

Information

  • Websitehuggingface.co
  • OrganizationsRopedia, Hugging Face
  • Published date2026/03/11

More Items

Hugging Face

Provides 315,000 pairwise human-preference votes comparing 15 English TTS models over 300 operational prompts, with 4,500 high‑quality audio renders and structured vote/pair/prompt records for training or evaluating preference/reward models. Metadata under CC-BY-4.0; audio use governed by model providers' terms.

Hugging Face

A 10‑billion‑document retrieval benchmark with per‑document 768‑dim unit‑norm dense embeddings and mGTE sparse embeddings, FineWeb text/metadata, and exact top‑1000 MS MARCO ground truth for ~120k queries. Built for large‑scale evaluation of dense/sparse/hybrid retrieval, filtered search, indexing, ANNS algorithms, and embedding compression.

Hugging Face

Provides 4.5 billion TikTok video records with captions, timestamps, music IDs and engagement counts for research; split across 27 zstd-compressed Parquet files (~289 GB) and sampled via TikTok's mobile API; released for research-use only with privacy and ToS caveats.