the-stack bigcode
Dataset Card for The Stack Changelog Release Description v1.0 Initial release of the Stack. Included 30 programming languages and 18 permissive licenses. Note: Three included licenses (MPL/EPL/LGPL) are considered weak copyleft licenses. The resulting near-deduplicated dataset is 3TB in size. v1.1 The three copyleft licenses ((MPL/EPL/LGPL) were excluded and the list of permissive licenses extended to 193 licenses in total. The list of programming… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack.
- 種別
- dataset
- ライセンス
- other
- 言語
- code
- ダウンロード
- 15,528
- いいね
- 1,058
- アクセス
- gated
- ファイル
- 0
タグ
- コードデータ
- 事前学習コーパス
- コード生成
- 大規模データセット
- 機械学習
- NLP
概要
The Stackはビッグコード社が公開した大規模なソースコードデータセットで、多数のプログラミング言語とパーミッシブライセンスで構成され、近重複排除後のサイズは約3TBに及びます。コピーレフトライセンスを除外し、ライセンス適法性を重視した設計で、ソフトウェア開発向け大規模言語モデルの事前学習・継続学習に適しています。コード生成やコード理解モデルの基盤データとして広く利用されています。