AIAny
Icon for item

Syn4D: A Multiview Synthetic 4D Dataset

Provides multiview synthetic RGB video clips with per-frame depth, instance masks, dense long-range 3D point tracks, camera poses, and SMPL‑X human pose/shape labels for 4D reconstruction, tracking, and geometry-aware novel-view synthesis. Includes ~4.7K clips (1.4M frames) and is licensed for AI training.

Introduction

Dense, multiview geometric supervision for dynamic scenes is scarce, and that scarcity restricts progress on 4D reconstruction, tracking, and geometry-aware novel-view synthesis. Syn4D supplies large-scale, fully synthetic multiview video with dense, queryable geometry so models can learn long-range spatiotemporal correspondences and geometry-consistent rendering.

What Sets It Apart
  • Dense, queryable 3D tracks: every pixel can be unprojected to 3D at any time and projected into any other camera/time via an efficient barycentric-map + mesh representation, enabling complete long-range 3D tracking benchmarks.
  • Multiview + geometry + humans: synchronized multi-camera RGB, per-pixel depth, camera trajectories, instance/mask metadata and SMPL‑X body pose/shape annotations in the same corpus — so tasks from pose estimation to 4D reconstruction can be trained jointly.
  • Scale and diversity: ~4.7K multiview clips (≈1.4M frames) rendered in Unreal Engine using a wide catalog of environments and 1,674 animated assets plus 585 simulated humans, curated to encourage generalization.
  • AI‑training license and tooling: all 3D assets are licensed for AI training, dataset packages include metadata mappings, visualizer code, and Kubric-style subsets to ease integration into training pipelines.
Who It's For and Trade-offs

Great fit if you need dense, geometry-consistent supervision for research or development in 4D reconstruction, multiview depth/camera estimation, long-range 3D tracking, or geometry-aware novel-view synthesis. Also useful for pretraining or augmenting human pose/SMPL‑X models when multiview geometry helps.

Look elsewhere if you require real-world photographic noise, raw sensor artifacts, or datasets for facial identity/facial-recognition tasks (human characters are synthetic and BEDLAM2-derived motion metadata is partially restricted). Be aware the full release is large (storage and decompression required) and synthetic-to-real domain gaps remain a consideration when transferring to real data.

Information

  • Websitehuggingface.co
  • OrganizationsVisual Geometry Group, University of Oxford, Nanyang Technological University, Naver Labs Europe
  • AuthorsZeren Jiang, Yushi Lan, Yihang Luo, Yufan Deng, Zihang Lai, Edgar Sucar, Christian Rupprecht, Iro Laina, Diane Larlus, Chuanxia Zheng
  • Published date2026/06/21

Categories

More Items

Hugging Face

Evaluates retrievers and search agents on synthetic multi-hop questions that require assembling a complete set of supporting evidence. Provides English and Russian variants (395 questions each), a fixed dense index embedded with Qwen3-Embedding-8B, and BrowseComp-Plus evaluation integrations.

Hugging Face

Provides re-annotated academic video instruction data for captioning, video QA, and fine-grained motion understanding; rewrites short answers and concise captions into evidence-grounded, instruction-following responses and supplies JSONL annotation files (original videos not included).

Hugging Face

Provides 324 Russian short-answer web-search tasks with gold supporting documents to evaluate fixed-index retrievers and search agents. Tasks span eight topical categories and five retrieval challenge types (multihop, structured evidence, temporal, entity disambiguation, comparative) and use a Qwen3-Embedding-8B index for evaluation.