Browser Use Bench v2.1 | 180 tasks | October 7, 2026

GPT-6 Lunaon real browser tasks.

#12 of 19 at xhigh effort. No other model in the test scores higher for less.

Score69.0
Cost / task16¢
Median time11 min
#ModelScoreCost / task
12GPT-6 Luna xhigh69.016¢
16GPT-6 Luna medium64.37.7¢

180 long tasks on live websites, the same for every model, run in the open-source BrowserCode harness (Browser Use's own agent runs in ours). Mean weighted rubric score per task, 0-100. Estimated model cost per task in USD. Excludes search, judge, browser and runner costs. Runner and encrypted tasks: github.com/browser-use/benchmark.