ShareGPT_Vicuna_unfiltered anon8231489123

Further cleaning done. Please look through the dataset and ensure that I didn't miss anything. Update: Confirmed working method for training the model: https://huggingface.co/AlekseyKorshuk/vicuna-7b/discussions/4#64346c08ef6d5abefe42c12c Two choices: Removes instances of "I'm sorry, but": https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/blob/main/ShareGPT_V3_unfiltered_cleaned_split_no_imsorry.json Has instances of "I'm sorry, but":… See the full description on the dataset page: https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered.

種別
dataset
ライセンス
apache-2.0
言語
en
ダウンロード
411,805
いいね
912
アクセス
public
ファイル
0

タグ

  • 対話データ
  • 英語テキストデータ
  • 事前学習
  • 指示微調整
  • 大規模言語モデル
  • データセット
  • NLP

概要

ShareGPTの約10万件のユーザー対話を英語のみに絞り、非英語・過剰なUnicode・AI道徳説教等の破棄文を除去して約5.3万件に精査した対話データセットです。Vicunaモデルの学習用に2048トークンのチャンクに分割済みで、フィルタリングなしの英語Vicunaモデルのファインチューニングに適しています。Apache-2.0ライセンスで公開されています。

README

--- license: apache-2.0 language: - en --- **追加のクリーニングを実施しました。データセットをご確認いただき、漏れがないかお確かめください。** **更新: モデルトレーニングの有効な手法を確認しました: https://huggingface.co/AlekseyKorshuk/vicuna-7b/discussions/4#64346c08ef6d5abefe42c12c** 選択肢は2つあります: - 「I'm sorry, but」のインスタンスを削除: https://huggingface.co/datasets/anon8231489…

查看完整页面 · 查看原文