stack-v3-train HuggingFaceCode

🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train.

種別
dataset
ライセンス
odc-by
言語
code
ダウンロード
237,737
いいね
344
アクセス
public
ファイル
500

タグ

  • コードデータ
  • 事前学習コーパス
  • コード生成
  • 多言語コーパス
  • データセット
  • 大規模言語モデル

概要

The Stack v3はGitHubから直接クロールされ、713言語、約173Mリポジトリからなる15.9TB(約4.9兆トークン)のソースコード事前学習データセットです。v3ではファイル内容がインラインで埋め込まれ、リポジトリ単位で整理されているため、ダウンロード後すぐにコードLLMの事前学習(特にリポジトリ文脈を考慮した学習)に利用できます。品質フィルタリングと近重複除去済みで、コードLLMのオープンで透明な事前学習を目的としています。

README

--- thumbnail: "https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train/resolve/main/assets/banner.png" annotations_creators: [] language_creators: - crowdsourced - expert-generated language: - code license: - odc-by multilinguality: - multilingual size_categories: - 100M<n<1B source_dataset…

查看完整页面 · 查看原文