BROWSER USE

- Browser Use Agents: give Browser Use a task and receive completed work. API V4 is current for new integrations.
- Browser Infrastructure: connect your agent or automation to managed browsers through SDK, REST, or CDP. Starts at $0.02/browser-hour.
- Developer tools: Open Source, Browser Harness, SDK, and MCP support the two products above.

[Developer Index](https://browser-use.com/index.md)
[Product Map](https://browser-use.com/llms.txt)
[Full Product Context](https://browser-use.com/llms-full.txt)
[Pricing](https://browser-use.com/pricing.md)
[Cloud Docs](https://docs.browser-use.com/cloud/quickstart)
[Open Source Docs](https://docs.browser-use.com/open-source/introduction)

---

# Benchmarks - Browser Use

These benchmarks compare Browser Use against other web automation frameworks and cloud browser providers on accuracy and stealth across real-world websites.

[Agent benchmarks](https://browser-use.com/benchmarks/agents)
[Browser benchmarks](https://browser-use.com/benchmarks/browsers)
[Cloud Platform](https://cloud.browser-use.com)

---

## Web Agent Benchmarks

### Internal Bench Hard

106 difficult tasks on live websites that a careful human can complete but most agents cannot. Every model runs on the same harness. A task counts as solved only when the final answer or end state is exactly right; partial credit earns nothing. Cost per solved task is total recorded spend divided by tasks solved. Updated August 1, 2026.

| Agent or model | Strict accuracy | Cost per solved task |
|---|---:|---:|
| Browser Use Agents | 82% | $0.17 |
| Opus 5 | 62% | $3.40 |
| Gemini 3.1 Pro | 59% | $2.20 |
| Sonnet 5 | 59% | $1.55 |
| GPT-5.6 | 52% | $1.10 |
| Gemini 3.6 Flash | 46% | $0.62 |
| GPT-5 | 37% | $0.44 |

[Current agent benchmark page](https://browser-use.com/benchmarks/agents)

### Online-Mind2Web

300 tasks across 136 live websites — shopping, finance, travel, government, and more. All 300 tasks used, no tasks removed. All tasks run on live websites. Scored by an agentic judge built on Claude Agent SDK, aligned with human judges. Testing conducted March 2026.

Benchmark source: https://github.com/OSU-NLP-Group/Online-Mind2Web
Blog post: https://browser-use.com/posts/online-mind2web-benchmark

| Agent | Accuracy |
|-------|----------|
| Browser Use Cloud (v4) | 98% |
| ABP + Opus 4.6 | 86% |
| TinyFish | 81% |
| Navigator | 78% |
| Gemini CUA | 69% |
| Stagehand (Gemini 2.5 CU) | 65% |
| OpenAI Operator | 61% |
| Sonnet 4.0 CU | 61% |
| Stagehand (Sonnet 4.5) | 55% |

### BU Bench V1

100 hand-selected tasks from WebBench, Mind2Web, GAIA, BrowseComp, and 20 custom page interaction challenges. All tasks are hard but verified completable. Each task run multiple times across different LLMs and agent settings. Scored by LLM judge (Gemini 2.5 Flash), 87% alignment with 200 hand-labeled traces. Testing conducted January 2026, updated March 2026.

Benchmark source: https://github.com/browser-use/benchmark
Blog post: https://browser-use.com/posts/ai-browser-agent-benchmark

| Model | Accuracy |
|-------|----------|
| Browser Use Cloud (bu-ultra) | 78.0% |
| OSS + ChatBrowserUse-2 | 63.3% |
| claude-opus-4-6 | 62.0% |
| gemini-3-1-pro | 59.3% |
| claude-sonnet-4-6 | 59.0% |
| gpt-5 | 52.4% |
| gpt-5-mini | 37.0% |
| gemini-2.5-flash | 35.2% |

---

## Stealth Benchmarks

### BrowserBench

Third-party benchmark created by Halluminate. 296 tasks across antibot-protected sites. Includes lower-security sites, so scores are generally higher than the BU Stealth Benchmark. Browser Use ran BrowserBench across all providers. Scoring is pass/fail based on whether the agent was blocked by antibot protection.

Benchmark source: https://github.com/Halluminate/browserbench
Blog post: https://browser-use.com/posts/stealth-benchmark

| Provider | Success Rate |
|----------|-------------|
| Browser Use Cloud | 84.8% |
| Hyperbrowser | 76.4% |
| Anchor | 76.0% |
| Steel | 73.3% |
| Browserbase | 70.3% |

### BU Stealth Benchmark

Built from 300,000 real production security check events. 71 high-security websites across Cloudflare, PerimeterX, Datadome, Akamai, reCaptcha, and others. Simple 3-step tasks per site — if it fails, the browser got blocked. Each provider tested multiple times with the same agent (bu-2-0) and model. Scored by LLM judge (gemini-2.5-flash) on whether the agent was blocked. Page load failures count as blocks. Controls: Headless Chromium scored 2%, Headful Chromium scored 50%.

Benchmark source: https://github.com/browser-use/benchmark
Blog post: https://browser-use.com/posts/stealth-benchmark

| Provider | Bypass Rate |
|----------|------------|
| Browser Use Cloud | 81% |
| Anchor | 77% |
| Onkernel | 67% |
| Steel | 47% |
| Browserbase | 42% |
| Hyperbrowser | 40% |
