databricks-dolly-15k databricks

Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/databricks/databricks-dolly-15k.

種別
dataset
ライセンス
cc-by-sa-3.0
言語
en
ダウンロード
43,318
いいね
1,075
アクセス
public
ファイル
0

タグ

  • 命令微調整
  • 対話データ
  • テキストデータ
  • 大規模言語モデル
  • 英語コーパス
  • 合成データ
  • 質問応答
  • NLP

概要

Databricks の従業員数千人が手作業で作成した、指示追従型の対話データセットです。InstructGPT 論文で示されたブレインストーミング、分類、閉域QA、生成、情報抽出、オープンQA、要約などのタスクカテゴリをカバーしています。大規模言語モデルの命令微調整に最適で、Self-Instruct 手法による合成データ生成やデータ拡張にも活用できます。商用利用も可能なオープンソースデータセットです。

README

--- license: cc-by-sa-3.0 task_categories: - question-answering - summarization language: - en size_categories: - 10K<n<100K --- # 概要 `databricks-dolly-15k` は、数千名の Databricks 従業員によって生成された、指示追従型レコードのオープンソースデータセットです。このデータセットは、[InstructGPT](https://arxiv.org/abs/2203.02155) 論文で概説されている複数の行動カテゴリ(ブレインストー…

查看完整页面 · 查看原文