the-stack bigcode

Dataset Card for The Stack Changelog Release Description v1.0 Initial release of the Stack. Included 30 programming languages and 18 permissive licenses. Note: Three included licenses (MPL/EPL/LGPL) are considered weak copyleft licenses. The resulting near-deduplicated dataset is 3TB in size. v1.1 The three copyleft licenses ((MPL/EPL/LGPL) were excluded and the list of permissive licenses extended to 193 licenses in total. The list of programming… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack.

类型
dataset
许可
other
语言
code
下载量
15,528
点赞
1,058
访问
gated
文件
0

标签

  • 代码数据
  • 预训练语料
  • 代码生成
  • 大语言模型
  • 多语言
  • 机器学习

摘要

The Stack 是 BigCode 组织发布的超大规模编程代码语料数据集,覆盖数百种编程语言与多种开源许可协议,经过近去重处理后体积达3TB。该数据集主要解决代码大语言模型的预训练数据来源问题,为代码生成模型(如 StarCoder)提供基础训练语料。适用于代码补全、代码生成、程序理解等任务的大模型预训练与微调。

查看完整页面 · 查看原文