fleurs google
FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.
- 类型
- dataset
- 许可
- cc-by-4.0
- 语言
- afr
- 下载量
- 97,349
- 点赞
- 443
- 访问
- public
- 文件
- 0
标签
- 语音识别
- 音频数据集
- 多语言语料
- 评测基准
- 语音
- 自动语音识别
摘要
FLEURS是FLoRes机器翻译基准的语音版本,也是XTREME-S评测基准的核心数据集,覆盖102种语言、10多个语系的语音识别任务。它提供约10小时监督训练数据,包含音频、转写、性别、语言标签等信息,用于评估多语言自动语音识别模型的跨语言泛化能力。适用于语音识别模型的训练与评测、多语言语音表示学习等场景。
README
--- annotations_creators: - 专家生成 - 众包 - 机器生成 language_creators: - 众包 - 专家生成 language: - afr - amh - ara - asm - ast - azj - bel - ben - bos - cat - ceb - cmn - ces - cym - dan - deu - ell - eng - spa - est - fas - ful - fin - tgl - fra - gle - glg - guj - hau - heb - hin - hrv - hun - hye - ind - i…