Alpaca-CoT QingyiSi

Instruction-Finetuning Dataset Collection (Alpaca-CoT) This repository will continuously collect various instruction tuning datasets. And we standardize different datasets into the same format, which can be directly loaded by the code of Alpaca model. We also have conducted empirical study on various instruction-tuning datasets based on the Alpaca model, as shown in https://github.com/PhoebusSi/alpaca-CoT. If you think this dataset collection is helpful to you, please like… See the full description on the dataset page: https://huggingface.co/datasets/QingyiSi/Alpaca-CoT.

类型
dataset
许可
apache-2.0
语言
en
下载量
3,875
点赞
785
访问
public
文件
0

标签

  • 指令微调
  • 对话数据
  • 代码数据
  • 多语言语料
  • 文本数据集
  • 大语言模型
  • NLP
  • 合成数据

摘要

Alpaca-CoT是一个持续收集各种指令微调数据集的合集,将不同来源的数据统一为标准格式,可直接被Alpaca模型代码加载。包含Alpaca、GPT-4生成、Chain-of-Thought推理、代码、金融、中文指令等多类数据,覆盖英语、中文等多语言,并基于Alpaca进行了实证研究。适用于大语言模型的指令微调与CoT推理能力训练。

README

--- language: - en - zh - ml tags: - Instruction - Cot license: apache-2.0 datasets: - dataset1 - dataset2 --- # 指令微调数据集合集(Alpaca-CoT) 本仓库将持续收集各类指令微调数据集。我们将不同的数据集统一为相同格式,可直接由 Alpaca 模型的[代码](https://github.com/PhoebusSi/alpaca-CoT)加载。 我们还基于 Alpaca 模型对各种指令微调数据集进行了实证研究,详情见 [https://github.com/PhoebusS…

查看完整页面 · 查看原文