ettin-pretraining-data jhu-clsp
Ettin Pre-training Data Phase 1 of 3: Diverse pre-training data mixture (1.7T tokens) used to train the Ettin model suite. This dataset contains the pre-training phase data used to train all Ettin encoder and decoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository. 📊 Data Composition Data Source Tokens (B) Percentage Description DCLM 837.2 49.1% High-quality web crawl data CC Head 356.6… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/ettin-pretraining-data.
- 类型
- dataset
- 许可
- mit
- 语言
- en
- 下载量
- 189,173
- 点赞
- 9
- 访问
- public
- 文件
- 0
标签
- 预训练语料
- 文本数据集
- 文本生成
- 多语言语料
- 代码数据
- 大语言模型
- 预训练
摘要
该数据集是Ettin编码器与解码器模型套件(17M-1B参数)训练阶段1的预训练数据混合物,包含1.7万亿token,涵盖网页抓取、代码、科学论文、数学内容等13个来源。数据以MDS格式提供,可直接配合Composer和ModernBERT训练仓库使用,支持流式加载。适用于大规模预训练语言模型的训练场景。
README
--- language: - en license: mit task_categories: - text-generation - fill-mask - text-classification - retrieval tags: - pretraining - language-modeling - encoder - decoder - foundation-model - transformer --- # Ettin 预训练数据 [](http…