openwebtext Skylion007
Dataset Card for "openwebtext" Dataset Summary An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2. This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances plain_text Size of downloaded dataset… See the full description on the dataset page: https://huggingface.co/datasets/Skylion007/openwebtext.
- Type
- dataset
- License
- cc0-1.0
- Language
- en
- Downloads
- 58,643
- Likes
- 530
- Access
- public
- Files
- 0
Tags
- 预训练语料
- 文本数据集
- 文本生成
- 语言建模
- 多语言语料
- NLP
README
--- annotations_creators: - no-annotation language_creators: - found language: - en license: - cc0-1.0 multilinguality: - monolingual pretty_name: OpenWebText size_categories: - 1M<n<10M source_datasets: - original task_categories: - text-generation - fill-mask task_ids: - language-modeling - maske…