hle cais

[!NOTE] IMPORTANT: Please help us protect the integrity of this benchmark by not publicly sharing, re-uploading, or distributing the dataset. Humanity's Last Exam 🌐 Website | 📄 Paper | GitHub Center for AI Safety & Scale AI Humanity's Last Exam (HLE) is a multi-modal benchmark at the frontier of human knowledge, designed to be the final closed-ended academic benchmark of its kind with broad subject coverage. Humanity's Last Exam consists of 2,500 questions across dozens… See the full description on the dataset page: https://huggingface.co/datasets/cais/hle.

类型
dataset
许可
mit
下载量
36,281
点赞
908
访问
gated
文件
0

标签

  • 评测基准
  • 多模态数据集
  • 大语言模型
  • 多模态
  • 数学推理
  • 问答

摘要

Humanity's Last Exam(HLE)是Center for AI Safety与Scale AI联合发布的多模态评测基准,由2500道覆盖数十个学科的前沿问题组成,旨在成为人类知识边界的最终封闭式学术基准。该数据集可评估大语言模型在多学科知识、数学推理与多模态理解上的极限能力,适用于评测前沿AI模型的能力上限与短板。

查看完整页面 · 查看原文