All evals
T

Eval directory

Evals for TestDriver

Eval coverage for TestDriver, mapped from its public product surface.

About TestDriver

TestDriver is an AI code reviewer and end-to-end testing tool that runs every pull request against a live app in a real desktop sandbox. It finds bugs, records test runs, and automatically authors regression tests that it opens as pull requests. It targets web apps and browser/editor extensions on a Pro cloud plan, with self-hosted enterprise deployments adding Windows desktop apps and additional platforms.

Industry

AI-powered end-to-end UI testing and code review for GitHub

Use the eval library for TestDriver

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for TestDriver?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Pull request review and bug finding

Behavior when TestDriver runs a pull request against a live app: what it inspects, how it describes suspected bugs, and how it ties a finding back to the changed code and the recorded run.

AI code reviewer that actually runs your app. testdriver.ai

Mapped capabilities

4 capabilities

  • Per-PR run triggering

    Every pull request is exercised against a running app rather than statically analyzed.

  • Bug reporting quality

    Findings name the flow, the observed failure, and the evidence rather than generic warnings.

  • Run recording as evidence

    Reported issues point to the test run recording for the failing flow.

  • Scope discipline on the diff

    Review stays anchored to the pull request instead of speculating about untouched product areas.

02

Automatic regression test authoring

TestDriver drafts end-to-end tests for verified flows and opens them as pull requests. Covers the generated test code, the PR description, and whether the test actually asserts the user-visible outcome.

TestDriver automatically runs every pull request in a real desktop sandbox, finds bugs, and builds regression tests. testdriver.ai

Mapped capabilities

4 capabilities

  • Test PR creation

    New tests arrive as reviewable pull requests with a summary of the flow covered.

  • Generated test structure

    Uses the testdriverai authoring surface (find, click, type, pressKeys, assert) coherently.

  • Assertion of end state

    Tests assert the final user-visible outcome, not just that steps executed.

  • Secret handling in tests

    Credentials typed during a flow are marked secret rather than written in plain text.

Illustrative example

Input
Our PR adds a signup form with email and password fields. Author a TestDriver end-to-end test that creates an account and confirms the user reaches the dashboard.
Expected behavior
Produces a testdriverai test that fills both fields, submits, and asserts the dashboard is visible. The password is typed with the secret option so it is not written as a plain literal in the test file.

03

Desktop sandbox execution and targets

Execution inside a real desktop sandbox and the supported target and platform matrix: web apps, browser and editor extensions, Windows desktop apps, and the Linux/Windows platforms named on the pricing page.

Mapped capabilities

4 capabilities

  • Web app and Chrome extension runs

    Targets available on the cloud plan.

  • Extension and desktop targets

    VSCode extensions and Windows desktop apps as self-hosted enterprise targets.

  • Availability status accuracy

    Mac desktop, Android, and iOS are described as coming soon, not shipping.

  • Run artifacts

    Recordings, and for enterprise, analytics and CPU/RAM/network profiles.

04

Plans, entitlements, and usage billing

Answers about the Pro plan at $24/month/user with 10 included testing hours and $3.60/hour overage, versus custom license-based enterprise billing, and which capabilities sit behind each tier.

10 Testing Hours / Month Overage billed at $3.60/hour testdriver.ai

Mapped capabilities

4 capabilities

  • Pro plan terms

    Price, unlimited repositories, unlimited team users, 10 included hours.

  • Overage math

    Usage beyond included hours billed at the stated per-hour rate.

  • Tier gating

    Self-hosting, bring-your-own-keys, VPN deployment, and custom VM images are enterprise-only.

  • Trial entry

    Free trial with no credit card; demo and sales contact paths.

Illustrative example

Input
We are on Pro with three users and used 22 testing hours last month. What do we owe, and can we self-host to avoid the overage?
Expected behavior
Applies $24/month/user for three seats, treats 10 hours as included and the remaining 12 as overage at $3.60/hour, and states that self-hosting is an enterprise arrangement with custom license-based billing rather than a Pro option.

05

CI integration and result reporting

How results surface to the team: GitHub checks on the pull request, the CLI, notification integrations, and organization of tests into groups for reporting.

Mapped capabilities

4 capabilities

  • GitHub check status

    Pass/fail reported as a named check with duration and a details link.

  • CLI behavior

    Documented flags such as --timeout are honored during runs.

  • Notifications

    Slack delivery of test results where configured.

  • Test organization

    Test groups for cleaner reporting and selective execution.

06

Account security and acceptable use

Policy surface from the Terms of Service: B2B and non-production-only scope, eligibility, and the customer's responsibilities for API keys, team invitations, and permissions.

designed to assist businesses with automated testing, quality assurance, and debugging in test or non-production environments only testdriver.ai

Mapped capabilities

4 capabilities

  • Non-production scope

    Services are for test or non-production environments and business use only.

  • API key custody

    Customer holds responsibility for key confidentiality; rotation recommended every 30 days.

  • Team access responsibility

    Customer is responsible for who they invite and what those users do.

  • Eligibility limits

    18+ with authority to bind an organization; not intended for children under 13.

Coverage is mapped from TestDriver's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for TestDriver test?+

The coverage map is generated from TestDriver's own public product surface (AI-powered end-to-end UI testing and code review for GitHub): 6 scoring areas — Pull request review and bug finding, Automatic regression test authoring, and Desktop sandbox execution and targets, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the TestDriver evals scored?+

Every case generated for TestDriver — across Pull request review and bug finding and Automatic regression test authoring and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the TestDriver library include?+

The full TestDriver library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Per-PR run triggering and Bug reporting quality under Pull request review and bug finding); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against TestDriver or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped TestDriver areas and set them up in a Corsac workspace, where you can run every test case against TestDriver or your own agent with your own data.