cosmopedia HuggingFaceTB

Cosmopedia v0.1 Image generated by DALL-E, the prompt was generated by Mixtral-8x7B-Instruct-v0.1 Note: Cosmopedia v0.2 is available at smollm-corpus User: What do you think "Cosmopedia" could mean? Hint: in our case it's not related to cosmology. Mixtral-8x7B-Instruct-v0.1: A possible meaning for "Cosmopedia" could be an encyclopedia or collection of information about different cultures, societies, and topics from around the world, emphasizing diversity and global… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/cosmopedia.

类型
dataset
许可
apache-2.0
语言
en
下载量
22,274
点赞
740
访问
public
文件
0

标签

  • 合成数据
  • 预训练语料
  • 文本数据集
  • 多教师
  • 大语言模型
  • 深度学习

摘要

Cosmopedia v0.1 是当前最大的开放合成数据集之一,包含超3000万文件、250亿token,由Mixtral-8x7B-Instruct模型基于RefinedWeb、RedPajama等网络语料作为种子样本生成合成教材、博客、故事和WikiHow文章。数据集按种子来源分为8个子集(如斯坦福教材、OpenStax、Khan Academy、网络样本等),并通过调整受众和提示风格提升多样性。主要面向合成数据研究与小型语言模型预训练,曾被用于训练Cosmo-1B模型。

README

--- dataset_info: - config_name: auto_math_text features: - name: prompt dtype: string - name: text_token_length dtype: int64 - name: text dtype: string - name: seed_data dtype: string - name: format dtype: string - name: audience dtype: string splits: - name: train num_bytes: 8777587297.907892 num…

查看完整页面 · 查看原文