ShareGPT_Vicuna_unfiltered anon8231489123
Further cleaning done. Please look through the dataset and ensure that I didn't miss anything. Update: Confirmed working method for training the model: https://huggingface.co/AlekseyKorshuk/vicuna-7b/discussions/4#64346c08ef6d5abefe42c12c Two choices: Removes instances of "I'm sorry, but": https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/blob/main/ShareGPT_V3_unfiltered_cleaned_split_no_imsorry.json Has instances of "I'm sorry, but":… See the full description on the dataset page: https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered.
- 種別
- dataset
- ライセンス
- apache-2.0
- 言語
- en
- ダウンロード
- 411,805
- いいね
- 912
- アクセス
- public
- ファイル
- 0
タグ
- 対話データ
- 英語テキストデータ
- 事前学習
- 指示微調整
- 大規模言語モデル
- データセット
- NLP
概要
ShareGPTの約10万件のユーザー対話を英語のみに絞り、非英語・過剰なUnicode・AI道徳説教等の破棄文を除去して約5.3万件に精査した対話データセットです。Vicunaモデルの学習用に2048トークンのチャンクに分割済みで、フィルタリングなしの英語Vicunaモデルのファインチューニングに適しています。Apache-2.0ライセンスで公開されています。
README
--- license: apache-2.0 language: - en --- **追加のクリーニングを実施しました。データセットをご確認いただき、漏れがないかお確かめください。** **更新: モデルトレーニングの有効な手法を確認しました: https://huggingface.co/AlekseyKorshuk/vicuna-7b/discussions/4#64346c08ef6d5abefe42c12c** 選択肢は2つあります: - 「I'm sorry, but」のインスタンスを削除: https://huggingface.co/datasets/anon8231489…