All evals
T

Eval directory

Evals for TestSprite

Eval coverage for TestSprite, mapped from its public product surface.

About TestSprite

TestSprite is an agentic testing platform that reads a team's code, designs, and PRDs, then exercises the running application like a real user to verify E2E, API, and visual-regression behavior. It returns machine-readable failure bundles (screenshots, DOM snapshots, root-cause hypotheses, suggested fixes) that coding agents can act on and re-run, and it lives in the terminal, CLI, IDE (via MCP), and CI. It is sold self-serve on a credit-based freemium plan with Free, Starter, Standard, and Enterprise tiers.

Industry

AI software testing / QA automation agent

Headquarters

Seattle, WA, United States

Use the eval library for TestSprite

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for TestSprite?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Context-grounded test generation

Turning a team's existing artifacts — code, PRDs, designs, and tickets — into executable test plans and test lists with minimal developer input, rather than requiring hand-written step-by-step instructions.

Reads your designs, code & tickets — integrates with the tools you already use. www.testsprite.com

Mapped capabilities

4 capabilities

  • PRD-to-test-plan derivation

    Generating executable test cases from product requirements, including requirements the code does not yet satisfy.

  • Codebase and API-doc grounding

    Using repository context and supplied API documentation to construct front-end and back-end workflow coverage.

  • Design and ticket integration

    Reading designs and tracker items from connected tools as test-generation context.

  • Edge-case and corner-case expansion

    Extending beyond happy paths to corner cases the marketing context claims are commonly missed.

02

Live application exercise

Actually driving the running product instead of reasoning about it statically: real browsers, live APIs, and no assertions against mocks, across E2E, API, and visual-regression modes.

Mapped capabilities

4 capabilities

  • URL-only autonomous exploration

    Pasting a live URL and having the platform explore and test it with no install.

  • Real-browser E2E execution

    Concurrent agents clicking through features as a real user would.

  • API and visual-regression modes

    Hitting live endpoints and detecting visual diffs alongside functional flows.

  • Session replay evidence

    Replaying any agent session as video for human inspection.

03

Failure bundles and the fix loop

What comes back when something breaks: one coherent, machine-readable bundle an agent can read, patch against, and re-run — the core differentiator versus flaky pass/fail reporting.

Mapped capabilities

4 capabilities

  • Bundle completeness

    Failing step plus surrounding steps, screenshot, DOM snapshot, and test source in a single artifact.

  • Root-cause hypothesis and suggested fix

    Actionable diagnosis rather than a bare failure signal.

  • Machine-readable verdicts

    Deterministic, parseable results instead of ambiguous pass/fail noise.

  • Fix-and-rerun closure

    Re-running the same test after an agent applies a change and reporting the delta.

Illustrative example

Input
Our checkout E2E test just failed on the payment step. Give me what my coding agent needs to fix it and re-run.
Expected behavior
Returns a single coherent bundle rather than a pass/fail line: the failing step with the steps around it, a screenshot, a DOM snapshot, the test source, a root-cause hypothesis, and a suggested fix, in a machine-readable form an agent can parse and act on.

04

Suite durability and regression memory

Keeping coverage trustworthy as the code changes: retaining passing tests, growing coverage per phase, and healing selector drift without masking genuine defects.

Mapped capabilities

4 capabilities

  • Passing-test retention across phases

    Accumulating coverage build over build rather than regenerating from scratch.

  • Regression detection on unrelated changes

    Catching previously-passing features broken by a single change.

  • Auto-Heal boundaries

    Repairing tests when the UI changes while still surfacing real bugs, not silently passing them.

  • Snapshot anchoring

    Binding a test to one snapshot rather than chasing a moving target.

05

Agent and developer surfaces

Meeting agents and developers where they already work — terminal, IDE, and CI — since a verifier for an overnight agent cannot assume a human watching a dashboard.

TestSprite uses your app like a real user — your agent fixes its own work before bugs reach you. www.testsprite.com

Mapped capabilities

4 capabilities

  • Open-source CLI

    Install and run via the published npm CLI package with an API key.

  • IDE plugin via MCP

    Driving test runs from Claude Code, Codex, and similar MCP clients.

  • GitHub Action and CI integration

    Autonomous testing attached to pull requests and CI pipelines.

  • Scheduled and unattended runs

    Test schedules that execute without an operator present, where the plan allows.

06

Plans, credits, and entitlements

Self-serve freemium commerce: credit quotas, test-list and schedule caps, model tiers, and Enterprise-only capabilities that must be stated accurately to prospects and to agents acting on their behalf.

TestSprite delivers the capabilities of a dedicated software test engineer, automating testing for both back-end and front-end systems. www.testsprite.com

Mapped capabilities

4 capabilities

  • Credit quota accuracy

    Monthly credit allowances and what consumes them per tier.

  • Test-list and schedule caps

    Limits that differ across Free, Starter, Standard, and Enterprise.

  • Model-tier differences

    Foundational versus advanced versus proprietary and custom-trained models by plan.

  • Enterprise-gated capabilities

    API access, custom configuration, and dedicated support boundaries.

Illustrative example

Input
I'm on the Free plan. How many credits and test lists do I get, and can I set up nightly scheduled regression runs?
Expected behavior
States 150 credits per month and 1 test list on Free, and says test schedules are not included at that tier — scheduling begins with Starter, which adds 5 test lists and 5 test schedules. Does not offer or imply a Free-tier scheduling workaround.

Coverage is mapped from TestSprite's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for TestSprite test?+

The coverage map is generated from TestSprite's own public product surface (AI software testing / QA automation agent): 6 scoring areas — Context-grounded test generation, Live application exercise, and Failure bundles and the fix loop, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the TestSprite evals scored?+

Every case generated for TestSprite — across Context-grounded test generation and Live application exercise and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the TestSprite library include?+

The full TestSprite library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, PRD-to-test-plan derivation and Codebase and API-doc grounding under Context-grounded test generation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against TestSprite or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped TestSprite areas and set them up in a Corsac workspace, where you can run every test case against TestSprite or your own agent with your own data.