the-stack bigcode
Dataset Card for The Stack Changelog Release Description v1.0 Initial release of the Stack. Included 30 programming languages and 18 permissive licenses. Note: Three included licenses (MPL/EPL/LGPL) are considered weak copyleft licenses. The resulting near-deduplicated dataset is 3TB in size. v1.1 The three copyleft licenses ((MPL/EPL/LGPL) were excluded and the list of permissive licenses extended to 193 licenses in total. The list of programming… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack.
- 类型
- dataset
- 许可
- other
- 语言
- code
- 下载量
- 15,528
- 点赞
- 1,058
- 访问
- gated
- 文件
- 0
标签
- 代码数据
- 预训练语料
- 代码生成
- 大语言模型
- 多语言
- 机器学习
摘要
The Stack 是 BigCode 组织发布的超大规模编程代码语料数据集,覆盖数百种编程语言与多种开源许可协议,经过近去重处理后体积达3TB。该数据集主要解决代码大语言模型的预训练数据来源问题,为代码生成模型(如 StarCoder)提供基础训练语料。适用于代码补全、代码生成、程序理解等任务的大模型预训练与微调。