finephrase HuggingFaceFW
Dataset Card for HuggingFaceFW/finephrase Dataset Summary Synthetic data generated by DataTrove: Model: HuggingFaceTB/SmolLM2-1.7B-Instruct (main) Source dataset: HuggingFaceFW/fineweb-edu, config sample-350BT, split train Generation config: temperature=1.0, top_p=1.0, top_k=50, max_tokens=2048, model_max_context=8192 Speculative decoding: {"method":"suffix","num_speculative_tokens":32} System prompt: None Input column: text Prompt families: faq prompt Rewrite… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finephrase.
- 类型
- dataset
- 许可
- odc-by
- 语言
- en
- 下载量
- 252,026
- 点赞
- 143
- 访问
- public
- 文件
- 0
标签
- 合成数据
- 预训练语料
- 指令微调
- 文本生成
- 问答
- 数据集
- NLP
- 大语言模型
摘要
FinePhrase是由Hugging Face使用SmolLM2-1.7B-Instruct模型基于FineWeb-Edu语料生成的超大规模合成文本数据集,包含约13.5亿样本。按FAQ、数学、表格、教程四种提示族重写文档,产出问答、数学应用题、表格问答和分步教程。主要用于大语言模型的预训练与指令微调,以提升模型的数学推理、表格理解和教学问答能力。
README
--- language: - en license: odc-by tags: - SmolLM2-1.7B-Instruct - fineweb-edu - synthetic - datatrove annotations_creators: - machine-generated language_creators: - found pretty_name: HuggingFaceFW/finephrase size_categories: - n>1M source_datasets: - HuggingFaceFW/fineweb-edu/sample-350BT task_ca…