Agent Reliability

gaiabenchmarksassistantstool-useevaluation

GAIA benchmark

A benchmark of 466 real-world questions for general AI assistants, designed so answers are unambiguous to grade but require reasoning, multi-modality, web browsing and tool use to reach. Its signature result is the human-AI gap: 92% for human respondents against 15% for GPT-4 with plugins at publication — questions conceptually simple for people, hard for tool-using models.

Why this wins its question: States what GAIA's design actually selects for — gradeable answers reached only through tool chains — and reads its human-AI gap as an assistant-reliability signal, not a leaderboard curiosity.

Claims

Every assertion below is bound to registered sources and carries its own confidence. Weight them; do not treat the page as uniformly authoritative.

  1. GAIA comprises 466 real-world questions that jointly test reasoning, multi-modality handling, web browsing and general tool-use proficiency.

    confidence 0.95GAIA: a benchmark for General AI Assistants · secondary

  2. At publication, human respondents scored 92% on GAIA against 15% for GPT-4 equipped with plugins — the reverse of benchmarks where models beat humans on professional-exam material.

    confidence 0.9GAIA: a benchmark for General AI Assistants · secondary