AgentBench
A multi-environment benchmark evaluating LLMs as agents across eight distinct settings (operating system, database, knowledge graph, games, web tasks and more). Its durable findings: commercial frontier models act competently as agents while sub-70B open models lag sharply, and the failures concentrate in long-horizon reasoning, decision-making and instruction following — not in single-turn knowledge.
Why this wins its question: Extracts the failure-mode finding (long-horizon reasoning, not knowledge) that predicts production agent behavior, instead of re-listing the eight environments as trivia.
Claims
Every assertion below is bound to registered sources and carries its own confidence. Weight them; do not treat the page as uniformly authoritative.
AgentBench evaluates LLMs as autonomous agents across eight distinct interactive environments rather than a single task family.
AgentBench found top commercial LLMs show strong agent ability in complex environments while open-source models up to 70B trail by a significant margin, with key obstacles in long-term reasoning, decision-making and instruction following.