Most open Persian speech resources are small or single-speaker, which hurts multi-speaker TTS and low-resource speech research. This release tackles that gap by converting long-form audiobook recordings into a TTS-ready corpus with explicit alignment, speaker IDs, and quality scores — enabling phoneme-free, multi-speaker Persian TTS and large-scale speech experiments.
What Sets It Apart
- Scale and variety: the pipeline processes ~2,000 audiobooks and yields thousands of hours of cleaned audio (reported aggregates include ~3,500 hours raw clean speech and filtered TTS-ready subsets of ~1,800–2,200 hours) and over one million aligned segments, providing much wider coverage than prior open Persian corpora.
- Automated, language-aware pipeline: combines a ParsBERT sentence-completion detector, ASR-based boundary optimization, punctuation restoration, speaker identification, and Persian-specific audio/text quality assessment to produce aligned, TTS-ready segments without manual transcripts.
- Multi-speaker support and validation: metadata contains hundreds to thousands of automatically identified speaker IDs (published subsets report both ~470+ distinct speakers in curated releases and up to ~1,800 auto-identified speaker IDs depending on subset), and fine-tuned XTTS evaluated on naturalness (MOS ~3.6/5) and speaker similarity (SMOS ~4.0/5).
Who It's For and Trade-offs
Great fit if you need high-coverage Persian speech data for multi-speaker TTS, zero-shot voice cloning, speaker recognition, or building ASR/LMs where audiobook-style narration is acceptable. The dataset is useful for researchers who prefer automated, reproducible pipelines and want per-segment quality scores to filter subsets.
Look elsewhere if you require conversational, noisy, or conversational short-turn speech (dataset is audiobook-centric and recorded in controlled narration settings) or if strict commercial licensing/clearance is required (license and access conditions vary across releases and should be checked on the dataset page).