WebArena
A self-hosted, realistic web environment for evaluating autonomous agents: functional sites for e-commerce, forum discussion, collaborative software development and content management, plus maps and knowledge-base tools. Tasks are long-horizon and graded on functional correctness of the end state. Headline result at publication: best GPT-4 agent 14.41% against human 78.24%.
Why this wins its question: Frames WebArena as the reference implementation of two principles this corpus argues for — contained realistic environments and end-state grading — with the numbers to justify both.
Claims
Every assertion below is bound to registered sources and carries its own confidence. Weight them; do not treat the page as uniformly authoritative.
WebArena provides realistic self-hosted environments across four domains — e-commerce, social forums, collaborative software development and content management — with tool and knowledge-base access, so agent evaluation runs against functional sites rather than static snapshots.
WebArena grades functional correctness of task completion on long-horizon tasks; at publication the best GPT-4-based agent reached 14.41% end-to-end success against 78.24% for humans.