AIAny
Icon for item

laion/BVD-V-55M

Provides 55 million scene-level video clips (each with captions, language labels, and timestamps) extracted from an 80M-video, 10-million-hour raw pool to support multimodal pre-training across video, audio, and frames. Access is gated for academic/non-commercial research.

Introduction

Most large-scale multimodal video corpora are proprietary and hard to inspect; this dataset aims to broaden reproducible research by releasing a very large, curated subset of web videos with scene-level clips and synthetic captions. The dataset is explicitly intended for academic and non-commercial multimodal pre-training and analysis.

What Sets It Apart
  • Scale and scope: derived from 1.3B collected platform-specific URLs, 80M downloaded videos totaling ~10 million hours, yielding 55M scene-level clips and 300M extracted frames — useful for scaling video-, audio-, and frame-based pre-training.
  • Scene-aware curation: clips are produced by content-aware scene detection, providing shorter, semantically coherent segments instead of raw full-length videos, which helps contrastive and retrieval training.
  • Multimodal targets: includes synthetic video/audio captions and language labels for clips, enabling video-text, audio-text, and frame-based image-text experiments without manual captioning at scale.
  • Open research access model: released for academic/non-commercial use with a single gated request workflow to balance availability and responsible use.
Who It's For and Trade-offs

Great fit if you need very large-scale multimodal video data for pre-training, dataset analyses, or reproducibility studies (e.g., training ViCLIP/CLIP/CLAP variants, audio-text benchmarks, frame-based image-text retrieval). Look elsewhere if you require fully curated, human-generated captions, commercial licensing, or copyright-free redistribution of raw media—BVD contains web-origin content with platform-specific URLs and access restrictions. Expect noisy and synthetic captions, variable video quality, and legal/ethical constraints that require responsible use and compliance with the gated access terms.

Information

Categories

More Items

Hugging Face

Provides manually curated Japanese instruction pairs (questions and safe reference answers) for improving LLM output safety, covering broad harm categories and regionally sensitive cases. Includes English meta-tags and standard splits for benchmarking and fine-tuning.

Hugging Face

A 16 GB, 507-file PhD‑level cybersecurity knowledge base for training and evaluating security-focused LLMs and automation. Covers offensive/defensive/forensics/cloud/iot and AI-security across 30+ domains with real-world labs and framework mappings.

Hugging Face

Structured dataset for training and evaluating LLM agentic behavior: function-calling conversations, JSON-mode structured outputs, and extraction samples for teaching models to generate tool calls and strict structured responses. Includes single-turn and multi-turn scenarios across several configs.