COIG BAAI

We propose the Chinese Open Instruction Generalist (COIG) project to maintain a harmless, helpful, and diverse set of Chinese instruction corpora. We welcome all researchers in the community to contribute to the corpus set and collaborate with us. We only release the first chip of COIG to help the Chinese LLMs' development in the exploration stage and appeal to more researchers joining us in building COIG. We introduce a manually verified translated general instruction corpus, a manually annotated exam instruction corpus, a human value alignment instruction corpus, a multi-round counterfactual correction chat corpus, and a leetcode instruction corpus. We provide these new instruction corpora to assist the community with instruction tuning on Chinese LLMs. These instruction corpora are also template workflows for how new Chinese instruction corpora can be built and expanded effectively.

类型
dataset
许可
apache-2.0
语言
zh
下载量
1,604
点赞
461
访问
public
文件
0

标签

  • 指令微调
  • 预训练语料
  • 多语言语料
  • 对话数据
  • 大语言模型
  • 文本数据集
  • NLP
  • 机器学习

摘要

COIG是BAAI推出的中文开放指令通用语料库,专为中文大语言模型的指令微调而构建。语料库涵盖人工校验的翻译指令、考试指令、人类价值对齐指令、反事实修正多轮对话和Leetcode代码指令五类数据,旨在提升中文LLM的指令遵循、推理和事实一致性能力。适用于中文大模型的指令跟随、推理与安全问题研究及微调训练。

README

--- license: apache-2.0 arxiv: 2304.07987 language: - zh --- # 中文开放指令通用项目 我们提出中文开放指令通用项目(**COIG**),旨在维护一套无害、有用且多样化的中文指令语料库。我们欢迎社区所有研究人员为语料库做出贡献并与我们合作。我们仅发布 COIG 的第一部分,以帮助中文大语言模型在探索阶段的发展,并呼吁更多研究人员加入我们共同建设 COIG。我们引入了人工验证的翻译通用指令语料、人工标注的考试指令语料、人类价值对齐指令语料、多轮反事实纠正对话语料以及 Leetcode 指令语料。我们提供这些新的指令语料,以协助社区对中…

查看完整页面 · 查看原文