dolphin QuixiAI
Dolphin 🐬 https://erichartford.com/dolphin Dataset details This dataset is an attempt to replicate the results of Microsoft's Orca Our dataset consists of: ~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl) ~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl) We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/dolphin.
- Type
- dataset
- License
- apache-2.0
- Language
- en
- Downloads
- 1,541
- Likes
- 434
- Access
- public
- Files
- 0
Tags
- 指令微调
- 文本生成
- 大语言模型
- 预训练语料
- 合成数据
- 文本数据集
README
--- license: apache-2.0 task_categories: - text-generation language: - en configs: - config_name: flan1m-alpaca-uncensored data_files: flan1m-alpaca-uncensored.jsonl - config_name: flan5m-alpaca-uncensored data_files: flan5m-alpaca-uncensored.jsonl --- Dolphin 🐬 https://erichartford.com/dolphin ##…