Browser Use Bench v2.1 | 180 tasks | October 7, 2026

Claude Haiku 5.5on real browser tasks.

#11 of 19 at xhigh effort. DeepSeek V4.1 Flash max scores 1.4 points higher for 2.6× less.

Haiku ran with a 32k output-token cap per request; GPT-6 Luna could use 128k. Both models support 128k.

Score69.5
Cost / task46¢
Median time14 min
#ModelScoreCost / task
11Claude Haiku 5.5 xhigh69.546¢
14Claude Haiku 5.5 max65.3$1.07
15Claude Haiku 5.5 high64.323¢
18Claude Haiku 5.5 medium53.49.2¢
19Claude Haiku 5.5 low45.62.7¢

180 long tasks on live websites, the same for every model, run in the open-source BrowserCode harness (Browser Use's own agent runs in ours). Mean weighted rubric score per task, 0-100. Estimated model cost per task in USD. Excludes search, judge, browser and runner costs. Runner and encrypted tasks: github.com/browser-use/benchmark.