Browser Use Bench v2.1 | 180 tasks | October 7, 2026

Claude Sonnet 5.5on real browser tasks.

#7 of 19. GPT-6.1 Sol medium scores 6.9 points higher for 1.2× less.

Score72.7
Cost / task$1.53
Median time9 min

180 long tasks on live websites, the same for every model, run in the open-source BrowserCode harness (Browser Use's own agent runs in ours). Mean weighted rubric score per task, 0-100. Estimated model cost per task in USD. Excludes search, judge, browser and runner costs. Runner and encrypted tasks: github.com/browser-use/benchmark.