AIAny
Icon for item

Masri 100h — Egyptian Arabic Single-Speaker Speech Corpus

100-hour, single-narrator Egyptian Arabic speech corpus with 15,653 aligned clips at 24 kHz for TTS and ASR fine-tuning; studio-consistent audio, machine-generated undiacritized transcripts, CC BY-NC 4.0 (research/non-commercial use).

Introduction

Why this matters

Large, studio-quality single-speaker collections in dialectal Arabic are rare. This corpus supplies ~100 hours of Egyptian Arabic narration from one consistent voice with native 24 kHz audio and aligned transcripts—enough data to fine-tune production TTS models rather than only perform few-shot cloning, and valuable for ASR adaptation to Egyptian (Masri) dialect.

What Sets It Apart
  • Single-speaker scale: ~100 hours (15,653 clips) recorded across 247 source episodes, concentrating long-form narrated utterances that support long-context prosody learning rather than conversational fragments.
  • Production consistency: studio narration on one microphone chain (consistent loudness, low noise), delivered as 24 kHz, 16-bit PCM WAV—no resampling required for modern neural vocoders.
  • Dialectal transcripts: machine-generated, normalized Egyptian Arabic (undiacritized, no punctuation), preserving colloquial orthography (e.g. عايز, بقى, ازاي) that differs from MSA corpora.
  • Rich metadata: per-clip ASR confidence, source video id, start/end offsets and a promo flag to filter sponsor segments; split assignment is disjoint by source video for robust evaluation.
  • Licensing & provenance: manifests/transcripts/segmentation released under CC BY-NC 4.0; underlying source recordings remain with original rights holders—dataset intended for research/non-commercial use.
Who It's For and Trade-offs

Great fit if you need to fine-tune TTS or adapt ASR to Egyptian Arabic using a single, consistent voice (audiobooks, IVR, conversational agents), or to study dialectal phonology and prosody with studio-quality narration. Use the provided confidence/promo fields and timestamps to build stricter training subsets.

Look elsewhere if you need multi-speaker, spontaneous conversational, noisy/telephony, or commercially licensed voice assets: transcripts are machine-generated (expect errors), text is undiacritized (models must learn vowelization), most clips are long-form (~18–19s), and the CC BY-NC 4.0 license prohibits commercial use without further permission. If a guaranteed single-speaker acoustic guarantee is required, run speaker-embedding verification and filter out outliers.

Information

  • Websitehuggingface.co
  • OrganizationsHugging Face
  • AuthorsEhab Negm
  • Published date2026/08/08

Categories

More Items

Hugging Face

Provides manually curated Japanese instruction pairs (questions and safe reference answers) for improving LLM output safety, covering broad harm categories and regionally sensitive cases. Includes English meta-tags and standard splits for benchmarking and fine-tuning.

Hugging Face

A 16 GB, 507-file PhD‑level cybersecurity knowledge base for training and evaluating security-focused LLMs and automation. Covers offensive/defensive/forensics/cloud/iot and AI-security across 30+ domains with real-world labs and framework mappings.

Hugging Face

Structured dataset for training and evaluating LLM agentic behavior: function-calling conversations, JSON-mode structured outputs, and extraction samples for teaching models to generate tool calls and strict structured responses. Includes single-turn and multi-turn scenarios across several configs.