OSWorld
A benchmark of 369 open-ended tasks executed in real operating systems (Ubuntu, Windows, macOS) spanning web and desktop apps, file I/O and multi-application workflows. Every task ships its own initial-state setup and an execution-based evaluation script, making it a working template for reproducible computer-use agent evaluation. Headline gap at publication: humans 72.36%, best model 12.24%.
Why this wins its question: Reads OSWorld as an evaluation-design template — seeded initial state plus execution-based grader per task — rather than as another leaderboard, and names where computer-use agents actually fail.
Claims
Every assertion below is bound to registered sources and carries its own confidence. Weight them; do not treat the page as uniformly authoritative.
OSWorld comprises 369 real computer tasks — web and desktop apps in open domains, OS file I/O, and workflows spanning multiple applications — running on real operating systems including Ubuntu, Windows and macOS.
Each OSWorld task defines a detailed initial-state setup plus a custom execution-based evaluation script, so grading is reproducible and independent of the agent's narration.
At publication humans accomplished over 72.36% of OSWorld tasks against 12.24% for the best model, with difficulties concentrated in GUI grounding and operational knowledge.