c4 allenai
C4 Dataset Summary A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset We prepared five variants of the data: en, en.noclean, en.noblocklist, realnewslike, and multilingual (mC4). For reference, these are the sizes of the variants: en: 305GB en.noclean: 2.3TB en.noblocklist: 380GB realnewslike: 15GB multilingual (mC4): 9.7TB (108 subsets, one… See the full description on the dataset page: https://huggingface.co/datasets/allenai/c4.
- 类型
- dataset
- 许可
- odc-by
- 语言
- af
- 下载量
- 1,430,934
- 点赞
- 630
- 访问
- public
- 文件
- 0
标签
- 预训练语料
- 多语言语料
- 文本数据集
- 大语言模型
- NLP
- 深度学习
摘要
C4 是 Google 对 Common Crawl 网页爬取语料的清洗版本,提供海量高质量文本用于大语言模型预训练,包含英语、多语言(mC4)、未清洗及新闻精选等多个变体。该数据集覆盖上百种语言,总规模达数 TB,是训练语言模型与掩码语言模型的经典基础语料,广泛用于学术研究与工业预训练。主要适用于语言建模、掩码建模等自监督预训练任务。
README
--- pretty_name: C4 annotations_creators: - no-annotation language_creators: - found language: - af - am - ar - az - be - bg - bn - ca - ceb - co - cs - cy - da - de - el - en - eo - es - et - eu - fa - fi - fil - fr - fy - ga - gd - gl - gu - ha - haw - he - hi - hmn - ht - hu - hy - id - ig - is …