mmlu cais

Dataset Card for MMLU Dataset Summary Measuring Massive Multitask Language Understanding by Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt (ICLR 2021). This is a massive multitask test consisting of multiple-choice questions from various branches of knowledge. The test spans subjects in the humanities, social sciences, hard sciences, and other areas that are important for some people to learn. This covers 57… See the full description on the dataset page: https://huggingface.co/datasets/cais/mmlu.

类型
dataset
许可
mit
语言
en
下载量
481,769
点赞
812
访问
public
文件
0

标签

  • 评测基准
  • 问答
  • 大语言模型
  • 多语言语料
  • 文本数据集

摘要

MMLU 是衡量大规模多任务语言理解能力的经典评测基准,包含 57 个学科领域约 1.4 万个多选选择题,覆盖人文、社科、理工等知识范畴。它主要用于评估和横向比较各类大语言模型在广泛学科知识上的问答推理能力,是检验 LLM 通用知识与理解水平的行业标准基准之一。

README

--- annotations_creators: - no-annotation language_creators: - expert-generated language: - en license: - mit multilinguality: - monolingual size_categories: - 10K<n<100K source_datasets: - original task_categories: - question-answering task_ids: - multiple-choice-qa paperswithcode_id: mmlu pretty_…

查看完整页面 · 查看原文