Skip to content
#

speech-dataset

Here are 40 public repositories matching this topic...

ManaTTS is the largest open Persian speech dataset with 114+ hours of transcribed audio. Includes data collection pipeline and tools. Suitable for Persian text-to-speech models.

  • Updated Jul 12, 2025
  • Jupyter Notebook

🇧🇮 The first large-scale, open-source speech and text dataset for Kirundi language. Building AI models for 12M+ Kirundi speakers through community collaboration. Includes ASR, TTS, and MT capabilities.

  • Updated May 7, 2026
  • Jupyter Notebook

Add this topic to your repo

To associate your repository with the speech-dataset topic, visit your repo's landing page and select "manage topics."

Learn more