the-stack bigcode

Dataset Card for The Stack Changelog Release Description v1.0 Initial release of the Stack. Included 30 programming languages and 18 permissive licenses. Note: Three included licenses (MPL/EPL/LGPL) are considered weak copyleft licenses. The resulting near-deduplicated dataset is 3TB in size. v1.1 The three copyleft licenses ((MPL/EPL/LGPL) were excluded and the list of permissive licenses extended to 193 licenses in total. The list of programming… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack.

種別
dataset
ライセンス
other
言語
code
ダウンロード
15,528
いいね
1,058
アクセス
gated
ファイル
0

タグ

  • コードデータ
  • 事前学習コーパス
  • コード生成
  • 大規模データセット
  • 機械学習
  • NLP

概要

The Stackはビッグコード社が公開した大規模なソースコードデータセットで、多数のプログラミング言語とパーミッシブライセンスで構成され、近重複排除後のサイズは約3TBに及びます。コピーレフトライセンスを除外し、ライセンス適法性を重視した設計で、ソフトウェア開発向け大規模言語モデルの事前学習・継続学習に適しています。コード生成やコード理解モデルの基盤データとして広く利用されています。

查看完整页面 · 查看原文