GAIA gaia-benchmark

GAIA dataset GAIA is a benchmark which aims at evaluating next-generation LLMs (LLMs with augmented capabilities due to added tooling, efficient prompting, access to search, etc). We added gating to prevent bots from scraping the dataset. Please do not reshare the validation or test set in a crawlable format. Data and leaderboard GAIA is made of more than 450 non-trivial question with an unambiguous answer, requiring different levels of tooling and autonomy to… See the full description on the dataset page: https://huggingface.co/datasets/gaia-benchmark/GAIA.

类型
dataset
语言
en
下载量
17,800
点赞
753
访问
gated
文件
0

标签

  • 评测基准
  • 大语言模型
  • Agent
  • 问答
  • 推理
  • 评测面板

摘要

GAIA是一个评估下一代大语言模型(结合工具调用、高效提示、搜索等增强能力)的基准数据集,包含450多个答案明确且需要不同层级工具调用与自主性的非平凡问题。它用来衡量模型在复杂现实任务中的工具使用与推理能力,并提供排行榜供横向对比,适用于大模型能力测评与Agent研究。

查看完整页面 · 查看原文