GAIA benchmark
A benchmark of 466 real-world questions for general AI assistants, designed so answers are unambiguous to grade but require reasoning, multi-modality, web browsing and tool use to reach. Its signature result is the human-AI gap: 92% for human respondents against 15% for GPT-4 with plugins at publication — questions conceptually simple for people, hard for tool-using models.
Why this wins its question: States what GAIA's design actually selects for — gradeable answers reached only through tool chains — and reads its human-AI gap as an assistant-reliability signal, not a leaderboard curiosity.
Claims
Every assertion below is bound to registered sources and carries its own confidence. Weight them; do not treat the page as uniformly authoritative.
GAIA comprises 466 real-world questions that jointly test reasoning, multi-modality handling, web browsing and general tool-use proficiency.
At publication, human respondents scored 92% on GAIA against 15% for GPT-4 equipped with plugins — the reverse of benchmarks where models beat humans on professional-exam material.