hle cais
[!NOTE] IMPORTANT: Please help us protect the integrity of this benchmark by not publicly sharing, re-uploading, or distributing the dataset. Humanity's Last Exam 🌐 Website | 📄 Paper | GitHub Center for AI Safety & Scale AI Humanity's Last Exam (HLE) is a multi-modal benchmark at the frontier of human knowledge, designed to be the final closed-ended academic benchmark of its kind with broad subject coverage. Humanity's Last Exam consists of 2,500 questions across dozens… See the full description on the dataset page: https://huggingface.co/datasets/cais/hle.
- 类型
- dataset
- 许可
- mit
- 下载量
- 36,281
- 点赞
- 908
- 访问
- gated
- 文件
- 0
标签
- 评测基准
- 多模态数据集
- 大语言模型
- 多模态
- 数学推理
- 问答
摘要
Humanity's Last Exam(HLE)是Center for AI Safety与Scale AI联合发布的多模态评测基准,由2500道覆盖数十个学科的前沿问题组成,旨在成为人类知识边界的最终封闭式学术基准。该数据集可评估大语言模型在多学科知识、数学推理与多模态理解上的极限能力,适用于评测前沿AI模型的能力上限与短板。