CulturaX uonlp
CulturaX Cleaned, Enormous, and Public: The Multilingual Fuel to Democratize Large Language Models for 167 Languages Dataset Summary We present CulturaX, a substantial multilingual dataset with 6.3 trillion tokens in 167 languages, tailored for large language model (LLM) development. Our dataset undergoes meticulous cleaning and deduplication through a rigorous pipeline of multiple stages to accomplish the best quality for model training, including language… See the full description on the dataset page: https://huggingface.co/datasets/uonlp/CulturaX.
- Type
- dataset
- Language
- af
- Downloads
- 12,768
- Likes
- 665
- Access
- gated
- Files
- 0
Tags
- 多语言语料
- 预训练语料
- 大语言模型
- 文本数据集
- NLP
- 数据处理