fineweb HuggingFaceFW

🍷 FineWeb 15 trillion tokens of the finest data the 🌐 web has to offer What is it? The 🍷 FineWeb dataset consists of more than 18.5T tokens (originally 15T tokens) of cleaned and deduplicated english web data from CommonCrawl. The data processing pipeline is optimized for LLM performance and ran on the 🏭 datatrove library, our large scale data processing library. 🍷 FineWeb was originally meant to be a fully open replication of 🦅 RefinedWeb, with a… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb.

種別
dataset
ライセンス
odc-by
言語
en
ダウンロード
418,229
いいね
3,211
アクセス
public
ファイル
500

タグ

  • 事前学習コーパス
  • テキスト生成
  • 大規模言語モデル
  • 英語語料
  • NLP
  • データセット

概要

FineWebは、CommonCrawlから抽出・精緻化された英語ウェブ文書を基にした、1兆トークン規模の大規模な事前学習用テキストデータセットです。品質フィルタリングと重複除去済みで、LLMの事前学習に最適化されており、公開コーパスとして最高水準の品質を誇ります。サンプル版(10BT/100BT/350BT)やCommonCrawl個別クロールごとの分割も提供され、多様な利用シーンに対応可能です。Hugging Faceの"Hallucinations Leaderboard"でも上位を占めるなど、LLM事前学習の標準的なデータ基盤として広く活用されています。

README

--- license: odc-by task_categories: - text-generation language: - en pretty_name: FineWeb size_categories: - n>1T configs: - config_name: default data_files: - split: train path: data/*/* - config_name: sample-10BT data_files: - split: train path: sample/10BT/* - config_name: sample-100BT data_fil…

查看完整页面 · 查看原文