BROWSER USE

- Browser Use Agents: give Browser Use a task and receive completed work. API V4 is current for new integrations.
- Browser Infrastructure: connect your agent or automation to managed browsers through SDK, REST, or CDP. Starts at $0.02/browser-hour.
- Developer tools: Open Source, Browser Harness, SDK, and MCP support the two products above.

[Developer Index](https://browser-use.com/index.md)
[Product Map](https://browser-use.com/llms.txt)
[Full Product Context](https://browser-use.com/llms-full.txt)
[Pricing](https://browser-use.com/pricing.md)
[Cloud Docs](https://docs.browser-use.com/cloud/quickstart)
[Open Source Docs](https://docs.browser-use.com/open-source/introduction)

---

# Claude Haiku 5.5 benchmark: browser agent score and cost

Browser Use Bench v2.1 | 180 tasks | October 7, 2026

## Claude Haiku 5.5on real browser tasks.

#11 of 19 at xhigh effort. DeepSeek V4.1 Flash max scores 1.4 points higher for 2.6× less.

Haiku ran with a 32k output-token cap per request; GPT-6 Luna could use 128k. Both models support 128k.

Score69.5

Cost / task46¢

Median time14 min

#

Model

Score

95% range

Cost / task

Median time

11

[Claude Haiku 5.5 xhigh](https://browser-use.com/benchmarks/models/claude-haiku-5-5)

69.5

65–73

46¢

14 min

14

[Claude Haiku 5.5 max](https://browser-use.com/benchmarks/models/claude-haiku-5-5)

65.3

61–70

$1.07

21 min

15

[Claude Haiku 5.5 high](https://browser-use.com/benchmarks/models/claude-haiku-5-5)

64.3

60–68

23¢

9 min

18

[Claude Haiku 5.5 medium](https://browser-use.com/benchmarks/models/claude-haiku-5-5)

53.4

50–57

9.2¢

5 min

19

[Claude Haiku 5.5 low](https://browser-use.com/benchmarks/models/claude-haiku-5-5)

45.6

42–50

2.7¢

2 min

180 long tasks on live websites, the same for every model, run in the open-source BrowserCode harness (Browser Use's own agent runs in ours). Mean weighted rubric score per task, 0-100. Estimated model cost per task in USD. Excludes search, judge, browser and runner costs. Runner and encrypted tasks: [github.com/browser-use/benchmark](https://github.com/browser-use/benchmark).

Compare |[full leaderboard](https://browser-use.com/benchmarks/agents)

-   [GPT-6.1 Sol79.6](https://browser-use.com/benchmarks/models/gpt-6-1-sol)
-   [GPT-6 Astra77.9](https://browser-use.com/benchmarks/models/gpt-6-astra)
-   [GPT-6 Sol74.4](https://browser-use.com/benchmarks/models/gpt-6-sol)
-   [GPT-5.6 Luna69.6](https://browser-use.com/benchmarks/models/gpt-5-6-luna)
-   [GPT-6 Luna69.0](https://browser-use.com/benchmarks/models/gpt-6-luna)
-   [Opus 5.575.0](https://browser-use.com/benchmarks/models/claude-opus-5-5)
-   [Opus 566.8](https://browser-use.com/benchmarks/models/claude-opus-5)
-   [Sonnet 5.572.7](https://browser-use.com/benchmarks/models/claude-sonnet-5-5)
-   [Fable 5.172.0](https://browser-use.com/benchmarks/models/claude-fable-5-1)
-   [Grok 4.773.3](https://browser-use.com/benchmarks/models/grok-4-7)
-   [DeepSeek V4.1 Flash70.9](https://browser-use.com/benchmarks/models/deepseek-v4-1-flash)
-   [GLM 5.3 FlashX59.7](https://browser-use.com/benchmarks/models/glm-5-3-flashx)

[Run your tasks in the cloud](https://cloud.browser-use.com/v4?utm_source=browser-use&utm_medium=benchmark-models)
