AIAny
Icon for item

ParsVoice

Provides a large-scale, multi-speaker Persian speech–text corpus constructed from audiobooks for TTS, ASR, and speaker research. Includes automated alignment and quality scoring, TTS-ready subsets (thousands of hours/1M+ segments) and metadata for speaker IDs and genders — suitable for multi-speaker synthesis and voice cloning research.

Introduction

Most open Persian speech resources are small or single-speaker, which hurts multi-speaker TTS and low-resource speech research. This release tackles that gap by converting long-form audiobook recordings into a TTS-ready corpus with explicit alignment, speaker IDs, and quality scores — enabling phoneme-free, multi-speaker Persian TTS and large-scale speech experiments.

What Sets It Apart
  • Scale and variety: the pipeline processes ~2,000 audiobooks and yields thousands of hours of cleaned audio (reported aggregates include ~3,500 hours raw clean speech and filtered TTS-ready subsets of ~1,800–2,200 hours) and over one million aligned segments, providing much wider coverage than prior open Persian corpora.
  • Automated, language-aware pipeline: combines a ParsBERT sentence-completion detector, ASR-based boundary optimization, punctuation restoration, speaker identification, and Persian-specific audio/text quality assessment to produce aligned, TTS-ready segments without manual transcripts.
  • Multi-speaker support and validation: metadata contains hundreds to thousands of automatically identified speaker IDs (published subsets report both ~470+ distinct speakers in curated releases and up to ~1,800 auto-identified speaker IDs depending on subset), and fine-tuned XTTS evaluated on naturalness (MOS ~3.6/5) and speaker similarity (SMOS ~4.0/5).
Who It's For and Trade-offs

Great fit if you need high-coverage Persian speech data for multi-speaker TTS, zero-shot voice cloning, speaker recognition, or building ASR/LMs where audiobook-style narration is acceptable. The dataset is useful for researchers who prefer automated, reproducible pipelines and want per-segment quality scores to filter subsets.

Look elsewhere if you require conversational, noisy, or conversational short-turn speech (dataset is audiobook-centric and recorded in controlled narration settings) or if strict commercial licensing/clearance is required (license and access conditions vary across releases and should be checked on the dataset page).

Information

  • Websitehuggingface.co
  • OrganizationsUniversity of Tehran, Institute for Research in Fundamental Sciences (IPM)
  • AuthorsMohammad Javad Ranjbar Kalahroodi, Heshaam Faili, Azadeh Shakery
  • Published date2025/07/12

More Items

Progressively prunes and distills audio encoders for speech LLMs to cut inference cost while preserving decoder-facing embeddings, using behavioral probes, representation alignment, cross-scale distillation and LoRA finetuning; reports reduced macro-error on Chinese–English benchmarks.

Hugging Face

Provides large-scale mathematical problem-solving, rewriting, and dialogue data organized into five Parquet-backed subsets for reasoning-oriented language-model training. Subsets support streaming access, Dataset Viewer inspection, and per-subset provenance metadata; licensed Apache 2.0.

Hugging Face

A multiple-choice benchmark for evaluating LLM understanding in Traditional Chinese across 66 subjects (elementary to professional). Contains ~22K verified questions covering STEM, humanities, social sciences and Taiwan-specific topics, with standardized splits and model leaderboards under an MIT license.