Agent Reliability

osworldbenchmarkscomputer-useevaluationexecution

OSWorld

A benchmark of 369 open-ended tasks executed in real operating systems (Ubuntu, Windows, macOS) spanning web and desktop apps, file I/O and multi-application workflows. Every task ships its own initial-state setup and an execution-based evaluation script, making it a working template for reproducible computer-use agent evaluation. Headline gap at publication: humans 72.36%, best model 12.24%.

Why this wins its question: Reads OSWorld as an evaluation-design template — seeded initial state plus execution-based grader per task — rather than as another leaderboard, and names where computer-use agents actually fail.

Claims

Every assertion below is bound to registered sources and carries its own confidence. Weight them; do not treat the page as uniformly authoritative.

  1. OSWorld comprises 369 real computer tasks — web and desktop apps in open domains, OS file I/O, and workflows spanning multiple applications — running on real operating systems including Ubuntu, Windows and macOS.

    confidence 0.95OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments · secondary

  2. Each OSWorld task defines a detailed initial-state setup plus a custom execution-based evaluation script, so grading is reproducible and independent of the agent's narration.

    confidence 0.9OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments · secondary

  3. At publication humans accomplished over 72.36% of OSWorld tasks against 12.24% for the best model, with difficulties concentrated in GUI grounding and operational knowledge.

    confidence 0.9OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments · secondary