PersonaHub proj-persona

Scaling Synthetic Data Creation with 1,000,000,000 Personas This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas: We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA HUB – a collection of 1 billion diverse personas automatically curated from web… See the full description on the dataset page: https://huggingface.co/datasets/proj-persona/PersonaHub.

Tipo
dataset
Licencia
cc-by-nc-sa-4.0
Lenguaje
en
Descargas
9,861
Me gusta
792
Acceso
public
Archivos
0

README

--- license: cc-by-nc-sa-4.0 task_categories: - text-generation - text-classification - token-classification - fill-mask - table-question-answering - text2text-generation language: - en - zh tags: - synthetic - text - math - reasoning - instruction - tool - persona size_categories: - 100M<n<1B conf…

查看完整页面 · 查看原文