01
Captcha Handling
Evaluates Browserbase's Captcha Handling across scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser infrastructure eval coverage.
Mapped capabilities
mapped

Eval directory · Code Assistant
Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Browserbase AI products.
Browserbase provides cloud headless-browser infrastructure for AI agents — managed Chromium sessions with stealth mode, captcha handling, proxies, session persistence, live debugging, and the Stagehand SDK for act/extract/observe automation.
Use the eval library for Browserbase
All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.
Related in Code Assistant
All evals →Coverage map
13 areas · 73 graded scenarios
Every eval set is graded on
Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls
01
Evaluates Browserbase's Captcha Handling across scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser infrastructure eval coverage.
Mapped capabilities
mapped
02
Evaluates Browserbase's Concurrency & Rate Limits across scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser infrastructure eval coverage.
Mapped capabilities
mapped
03
Evaluates Browserbase's Live Debugging & Session Inspector across scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser infrastructure eval coverage.
Mapped capabilities
mapped
04
Evaluates Browserbase's Proxy & Geo Routing across scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser infrastructure eval coverage.
Mapped capabilities
mapped
05
Evaluates Browserbase's Stealth & Fingerprinting across scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser infrastructure eval coverage.
Mapped capabilities
mapped
06
Evaluates Browserbase's Auth & Concurrency across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser Infrastructure for AI Agents eval coverage.
Mapped capabilities
9 scenarios
Public sample case
07
Evaluates Browserbase's Live View / Debug & Recordings across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser Infrastructure for AI Agents eval coverage.
Mapped capabilities
9 scenarios
Public sample case
08
Evaluates Browserbase's Page Automation Primitives across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser Infrastructure for AI Agents eval coverage.
Mapped capabilities
9 scenarios
Public sample case
09
Evaluates Browserbase's Safety, Consent & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser Infrastructure for AI Agents eval coverage.
Mapped capabilities
10 scenarios
10
Evaluates Browserbase's Session Lifecycle across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser Infrastructure for AI Agents eval coverage.
Mapped capabilities
9 scenarios
11
Evaluates Browserbase's Session Persistence & Recovery across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser Infrastructure for AI Agents eval coverage.
Mapped capabilities
9 scenarios
12
Evaluates Browserbase's act / extract / observe across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser Infrastructure for AI Agents eval coverage.
Mapped capabilities
9 scenarios
13
Evaluates Browserbase's Stealth & Anti-bot across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser Infrastructure for AI Agents eval coverage.
Mapped capabilities
9 scenarios
Each eval pack tests Browserbase's public product surface — including Captcha Handling, Concurrency Rate Limits, and Live Debugging Session Inspector — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.
Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Browserbase cases — from Safety Consent And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Browserbase library.
The Browserbase eval library includes 73 graded test cases across 13 eval packs, the largest being Safety Consent And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.
Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 13 Browserbase packs — Captcha Handling and Concurrency Rate Limits and the rest — against Browserbase or your own agent with your own data.