Emilia-Dataset amphion

Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline. News 🔥 2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.

Type
dataset
License
cc-by-4.0
Language
zh
Downloads
51,576
Likes
477
Access
gated
Files
0

Tags

  • 音频数据集
  • 语音合成
  • 音频数据处理
  • 语音
  • 多语言
  • 预训练语料
  • 文本生成语音

查看完整页面 · 查看原文