the-stack-v2 bigcode

The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset <-- you are here bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2.

类型
dataset
许可
other
语言
code
下载量
12,690
点赞
615
访问
gated
文件
0

标签

  • 代码数据
  • 预训练语料
  • 代码生成
  • 大语言模型
  • 多语言语料
  • 文本数据集

摘要

The Stack v2 是 BigCode 推出的大型多语言代码预训练语料数据集,覆盖600多种编程语言,并提供多个去重与过滤版本(如 dedup、train-full-ids、train-smol-ids),满足不同规模的训练需求。该数据集主要服务于大语言模型的代码生成能力训练,是代码大模型与代码智能领域的重要基础资源。

查看完整页面 · 查看原文