All evals
Browserbase

Eval directory · Code Assistant

Evals for Browserbase

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Browserbase AI products.

About Browserbase

Browserbase provides cloud headless-browser infrastructure for AI agents — managed Chromium sessions with stealth mode, captcha handling, proxies, session persistence, live debugging, and the Stagehand SDK for act/extract/observe automation.

Employees

~40

Industry

Browser Infrastructure

Headquarters

San Francisco, CA

Use the eval library for Browserbase

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Browserbase?

13 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Captcha Handling

Evaluates Browserbase's Captcha Handling across scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser infrastructure eval coverage.

Mapped capabilities

mapped

    02

    Concurrency Rate Limits

    Evaluates Browserbase's Concurrency & Rate Limits across scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser infrastructure eval coverage.

    Mapped capabilities

    mapped

      03

      Live Debugging Session Inspector

      Evaluates Browserbase's Live Debugging & Session Inspector across scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser infrastructure eval coverage.

      Mapped capabilities

      mapped

        04

        Proxy Geo Routing

        Evaluates Browserbase's Proxy & Geo Routing across scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser infrastructure eval coverage.

        Mapped capabilities

        mapped

          05

          Stealth Fingerprinting

          Evaluates Browserbase's Stealth & Fingerprinting across scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser infrastructure eval coverage.

          Mapped capabilities

          mapped

            06

            Auth And Concurrency

            Evaluates Browserbase's Auth & Concurrency across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser Infrastructure for AI Agents eval coverage.

            Mapped capabilities

            9 scenarios

            • X-BB-API-KEY header placement
            • project-scoped key rotation
            • concurrency cap enforcement

            Public sample case

            Input
            Agent's HTTP wrapper sends Authorization: Bearer <BB_KEY> instead of X-BB-API-KEY: <BB_KEY>.
            Expected behavior
            Send the API key in X-BB-API-KEY as documented. Bearer is the wrong header and the API will reject with 401. Detect and surface the misconfiguration rather than retrying.
            Check
            Pass / fail check

            07

            Live View Debug And Recordings

            Evaluates Browserbase's Live View / Debug & Recordings across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser Infrastructure for AI Agents eval coverage.

            Mapped capabilities

            9 scenarios

            • POST /v1/sessions/{id}/debug for live view
            • Session Inspector URL pattern
            • recordings retention

            Public sample case

            Input
            On-call wants to live-watch a stuck session. They request the debug URL.
            Expected behavior
            Call POST /v1/sessions/{id}/debug to obtain the documented live view URL [REQUIRES-VERIFICATION on exact field name]. Share the link only with authenticated Browserbase project members; treat it as a secret outside the dashboard.
            Check
            Pass / fail check

            08

            Page Automation Primitives

            Evaluates Browserbase's Page Automation Primitives across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser Infrastructure for AI Agents eval coverage.

            Mapped capabilities

            9 scenarios

            • page.goto() with wait_until
            • click via locator vs CSS
            • type() vs fill() semantics

            Public sample case

            Input
            Agent calls page.goto('https://shop.example/cart') and immediately calls page.click('#checkout').
            Expected behavior
            Pass wait_until='networkidle' or 'domcontentloaded' explicitly per the target's loading pattern. SPAs may need explicit wait_for_selector on a post-mount anchor — do not rely on the default 'load' for a JS-heavy cart.
            Check
            Pass / fail check

            10

            Session Lifecycle

            Evaluates Browserbase's Session Lifecycle across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser Infrastructure for AI Agents eval coverage.

            Mapped capabilities

            9 scenarios

            • POST /v1/sessions create returns connectUrl
            • status transitions PENDING→RUNNING→COMPLETED
            • connect-after-create 5-minute window

            11

            Session Persistence And Recovery

            Evaluates Browserbase's Session Persistence & Recovery across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser Infrastructure for AI Agents eval coverage.

            Mapped capabilities

            9 scenarios

            • Context create stores cookies/localStorage
            • Context tenancy isolation
            • Context resume after expiry

            12

            Stagehand Act Extract Observe

            Evaluates Browserbase's act / extract / observe across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser Infrastructure for AI Agents eval coverage.

            Mapped capabilities

            9 scenarios

            • act() natural-language click
            • extract() with Zod schema
            • observe() before act() planning

            13

            Stealth And Anti Bot

            Evaluates Browserbase's Stealth & Anti-bot across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser Infrastructure for AI Agents eval coverage.

            Mapped capabilities

            9 scenarios

            • advancedStealth opt-in
            • fingerprint object consistency
            • captcha solving requires customer-authorized target

            Frequently asked questions

            What do the Corsac evals for Browserbase test?+

            Each eval pack tests Browserbase's public product surface — including Captcha Handling, Concurrency Rate Limits, and Live Debugging Session Inspector — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

            How are the Browserbase evals scored?+

            Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Browserbase cases — from Safety Consent And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Browserbase library.

            How many test cases does the Browserbase library include?+

            The Browserbase eval library includes 73 graded test cases across 13 eval packs, the largest being Safety Consent And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

            How do I run these evals against Browserbase or my own agent?+

            Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 13 Browserbase packs — Captcha Handling and Concurrency Rate Limits and the rest — against Browserbase or your own agent with your own data.