Emilia-Dataset amphion
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline. News 🔥 2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.
- 类型
- dataset
- 许可
- cc-by-4.0
- 语言
- zh
- 下载量
- 51,576
- 点赞
- 477
- 访问
- gated
- 文件
- 0
标签
- 音频数据集
- 语音合成
- 音频数据处理
- 语音
- 多语言
- 预训练语料
- 文本生成语音
摘要
Emilia 是一个面向大规模语音生成的多语言多样化语音数据集,提供超20万小时(Emilia-Large)的高质量语音数据,涵盖多种语言与说话场景。该项目还包含 Emilia-Pipe 语音数据预处理管线源码,用于系统化清洗和构建高质量语音训练语料。适用于大模型语音合成、语音克隆与文本生成语音等语音生成任务的预训练与微调。