pile EleutherAI
The Pile is a 825 GiB diverse, open source language modelling data set that consists of 22 smaller, high-quality datasets combined together.
- 유형
- dataset
- 라이선스
- other
- 언어
- en
- 다운로드
- 2,877
- 좋아요
- 502
- 접근
- public
- 파일
- 0
README
--- annotations_creators: - no-annotation language_creators: - found language: - en license: other multilinguality: - monolingual pretty_name: the Pile size_categories: - 100B<n<1T source_datasets: - original task_categories: - text-generation - fill-mask task_ids: - language-modeling - masked-lang…