All evals
C

Eval directory

Evals for Checksum

Eval coverage for Checksum, mapped from its public product surface.

About Checksum

Checksum is a continuous quality platform whose AI agents autonomously generate, run, and heal end-to-end, API, and CI tests for web applications. Tests are delivered as standard Playwright code into the customer's repository via pull requests, and run in CI on every commit or PR. It is sold as a 'Results as a Service' offering priced by the number of maintained workflows, with human engineer verification of delivered tests.

Industry

AI software testing / QA automation platform

Use the eval library for Checksum

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Checksum?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Autonomous E2E Test Generation

The background E2E agent detects what matters in the application and produces production-ready Playwright tests without prompt-by-prompt supervision, bootstrapping a first suite in the initial week.

Bootstrap to 100-150 tests in your first week. checksum.ai

Mapped capabilities

4 capabilities

  • Flow detection and prioritization

    Analyzing the connected application to identify the most important user workflows to cover, including flows nobody explicitly demoed.

  • Playwright test authoring

    Emitting production-ready, standard Playwright test code rather than proprietary scripts or recorded fixtures.

  • Bootstrap coverage volume

    Reaching an initial suite on the order of 100-150 workflows in the first week and scaling from zero to thousands of tests.

  • Environment and setup configuration

    Connecting the repository and configuring the testing environment, credentials, and target app during onboarding.

02

Change-Scoped CI Testing

The CI agent runs on every commit, PR, and deployment, generating and actually executing tests targeted at the code that changed so results are ready before human review.

CI Guard generates 50-200 tests per PR, testing exactly what changed. checksum.ai

Mapped capabilities

4 capabilities

  • Diff-scoped test targeting

    Selecting and generating tests that exercise the exact code changed in a pull request rather than rerunning or regenerating everything.

  • Per-PR generation volume

    Producing 50-200 tests per pull request as described for CI Guard.

  • Pre-review execution and verification

    Executing generated tests and attaching verified results so the PR arrives already run, not merely proposed.

  • Continuous CI/CD triggering

    Staying plugged into CI to run on every commit and deployment so feedback arrives while changes are fresh.

Illustrative example

Input
A pull request modifies only the checkout coupon-code component. Ask Checksum's CI agent to cover this PR before review.
Expected behavior
Checksum generates tests exercising the coupon-code paths touched by the diff, executes them in CI, and reports concrete pass/fail results on the PR. It does not regenerate or rerun the unrelated remainder of the suite.

03

API Coverage at Scale

The API agent covers large endpoint surfaces in days rather than months, working alongside E2E and unit tests for combined protection.

Mapped capabilities

3 capabilities

  • Endpoint discovery

    Enumerating the API surface so that thousands of endpoints can be brought under coverage.

  • API test generation throughput

    Generating coverage at a volume matched to AI-accelerated output, on the order of thousands of tests.

  • Layered coverage composition

    E2E, API, and unit tests working together rather than as disconnected suites.

04

Test Durability and Recovery

Delivered suites are expected to be zero-maintenance: tests heal when the app evolves, failing runs are recovered in real time, and production incidents feed back into coverage.

Tests delivered as standard Playwright code to your repo — you own them checksum.ai

Mapped capabilities

4 capabilities

  • Auto-healing on app change

    Automatically repairing tests broken by product changes so the customer does not edit or maintain them.

  • CLI auto-recovery

    The Checksum CLI attempting to fix failing tests in real time during a run, distinct from post-hoc auto-healing.

  • Flaky-test reduction

    Reducing flakiness so suite results stay trustworthy, a theme Checksum addresses directly.

  • Production errors become tests

    Monitoring production errors and converting each surfaced bug into a maintained test.

Illustrative example

Input
A maintained login workflow test fails after the submit button's data-testid is renamed from 'login-submit' to 'auth-submit' in the app.
Expected behavior
Checksum heals the test on its own and opens a pull request with updated Playwright code using the new selector. The test's assertions and covered workflow stay unchanged, and the healed test passes when run.

05

Delivery, Ownership, and Repo Integration

Output lands in the customer's own repository as pull requests of pure Playwright code, with human engineer verification before delivery and no vendor lock-in.

Pure Playwright test code delivered to your repository as pull requests checksum.ai

Mapped capabilities

4 capabilities

  • Pull-request delivery

    Tests delivered into the customer repo as PRs with inline comments, webhooks, and auto-merge support.

  • Repository platform support

    GitHub App and native GitLab integration, including GitHub Enterprise and self-hosted GitLab instances.

  • Portability and ownership

    Tests are pure Playwright the customer owns and can run anywhere, including with Playwright directly.

  • Human engineer verification

    Final verification by a Checksum engineer so delivered tests are ready-to-go rather than unvetted AI output.

06

Reporting, Notifications, and Commercial Model

Run results flow to a dashboard and to team channels through a configurable integration layer, under pricing tied to the number of maintained workflows.

Mapped capabilities

4 capabilities

  • Dashboard reporting

    CLI runs auto-uploading reports to the dashboard for local and CI executions.

  • Webhook delivery

    Custom HTTP webhooks with HMAC-SHA256 signing, 48+ event types, retry logic, and flexible authentication.

  • Team notification channels

    Slack, Microsoft Teams, Discord, Google Chat, and transactional email notifications for results and bug alerts.

  • Workflow-based entitlements

    Pricing by maintained workflows with no per-seat or per-run charges, across Emerging/Scaling tiers and a 30-day trial.

Coverage is mapped from Checksum's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Checksum test?+

The coverage map is generated from Checksum's own public product surface (AI software testing / QA automation platform): 6 scoring areas — Autonomous E2E Test Generation, Change-Scoped CI Testing, and API Coverage at Scale, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Checksum evals scored?+

Every case generated for Checksum — across Autonomous E2E Test Generation and Change-Scoped CI Testing and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Checksum library include?+

The full Checksum library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Flow detection and prioritization and Playwright test authoring under Autonomous E2E Test Generation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Checksum or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Checksum areas and set them up in a Corsac workspace, where you can run every test case against Checksum or your own agent with your own data.