BROWSER USE DEVELOPER RESOURCES

- Hosted Agents (API V4): submit a task and receive the result from a managed browser agent.
- Browser Infrastructure: control managed remote browsers through SDK, REST, or CDP. Starts at $0.02/browser-hour.
- Open Source Library: build and run browser agents with the Python library.

[Developer Index](https://browser-use.com/index.md)
[Product Map](https://browser-use.com/llms.txt)
[Pricing](https://browser-use.com/pricing.md)
[Cloud Docs](https://docs.browser-use.com/cloud/quickstart)
[Open Source Docs](https://docs.browser-use.com/open-source/introduction)

---

# Web Agent Benchmarks - Browser Use

Every task we run, against every model, on the same harness. Including the runs we lose.

### Internal Bench Hard

Browser Use: 82% of tasks solved at 17¢ each. That is 20 points better than Opus 5, which costs 20× more per solved task.

cheaper and better than every model← cheaperbetter ↑0¢$1.00$2.00$3.0020%30%40%50%60%70%80%90%strict accuracycost per solved taskBrowser UseGPT-5Gemini 3.6 FlashGPT-5.6Sonnet 5Gemini 3.1 ProOpus 5

Toolstrict accuracyCost

Browser Use82%$0.17

Opus 562%$3.40

Gemini 3.1 Pro59%$2.20

Sonnet 559%$1.55

GPT-5.652%$1.10

Gemini 3.6 Flash46%$0.62

GPT-537%$0.44

Internal Bench Hard · updated 2026-08-01 · [all benchmarks](https://browser-use.com/benchmarks/agents)

What it measures

Our hardest internal set: 106 tasks that a careful human can finish but most agents cannot. Every model runs on the same harness, and we count a task solved only on a strict match, so partial credit earns nothing.

How it was run

-   Tasks: 106 hard tasks on live websites, none removed.
-   Cost: total recorded spend for the run divided by the number of tasks actually solved.
-   Scoring: strict — a task counts only when the final answer or end state is exactly right.
