All evals
QA Wolf

Eval directory

Evals for QA Wolf

Eval coverage for QA Wolf, mapped from its public product surface.

About QA Wolf

QA Wolf is an AI testing platform that maps application workflows, automates web and mobile end-to-end tests, and runs them fully in parallel. It is offered both as a self-serve platform priced per AI credit and runner minute, and as a managed "Coverage as a Service" offering with guaranteed test coverage. The service covers web, iOS, Android, and Electron apps, exports open-source Playwright code, and includes human-verified bug reports and 24-hour maintenance.

Industry

AI end-to-end software testing platform

Use the eval library for QA Wolf

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for QA Wolf?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Workflow Discovery & App Mapping

How the product builds and maintains a map of what there is to test, combining autonomous exploration with domain knowledge supplied by the customer team.

“AI autonomously explores your app and documents its workflows, getting domain knowledge from you to fill in any gaps.” www.qawolf.com

Mapped capabilities

4 capabilities

  • Autonomous app exploration

    AI explores the application and documents the workflows it finds.

  • Domain-knowledge gap filling

    Solicits knowledge from the customer team where exploration leaves gaps.

  • Existing test-plan review

    Pass over manual test plans to surface coverage gaps, redundancies, and untested edge cases.

  • Incident and spec grounding

    Uses past P0s/incident reports and roadmaps/specs to ground what coverage is needed.

02

Test Automation Authoring

How test cases become production-grade code, including hard cases the marketing surface calls out explicitly.

“The Automation AI writes production-grade code for complex web, iOS, and Android test cases.” www.qawolf.com

Mapped capabilities

4 capabilities

  • Bulk test automation

    Automating many cases from the mapped workflow set rather than one at a time.

  • Complex-capability cases

    Named hard cases: Canvas API, iBeacon detection, barcode scanning, map zone creation.

  • Data setup and teardown

    Test data provisioned and cleaned up within each test.

  • Maintenance for UI changes

    Updating existing tests when the application UI shifts.

03

Parallel Execution & Orchestration

How runs are scheduled, isolated, and coordinated so that a full suite completes without tests interfering with one another.

Mapped capabilities

4 capabilities

  • Unlimited parallel runs

    Whole suite executed concurrently rather than serially.

  • Per-test containerization

    Individually containerized tests with safeties against collisions and race conditions.

  • Run triggering

    Manual, scheduled, or deploy-time via webhook; pre-merge PR-branch suites.

  • Multi-user, multi-device coordination

    Run Rules-style orchestration across devices, services, and time.

04

Failure Triage & Maintenance

What happens after a test fails — the investigation loop that is meant to convert raw failures into a clean release signal.

“Guaranteed zero flakes. You will never be alerted to a test flake, only real bugs get flagged.” www.qawolf.com

Mapped capabilities

4 capabilities

  • Human failure investigation

    Every failure investigated, with 24-hour maintenance and repair of the test if needed.

  • Zero-flake filtering

    Flakes absorbed rather than alerted; only verified issues reach the team.

  • Human-verified bug reports

    Report contents: video with repro steps, Playwright traces, console logs.

  • Release-signal clarity

    Noise-free results the customer team is not expected to manage.

Illustrative example

Input
One test failed in last night's run. Will that be sent to my team as a bug, and what will the report contain?
Expected behavior
Explains that the failure is investigated and reproduced by a human before anything is flagged, so flakes are filtered out and only verified bugs reach the team, and that the report includes a video with reproduction steps, Playwright traces, and console logs.

05

Platform Coverage & Portability

The breadth of application types and browsers supported, and how easily the resulting assets and results move into the customer's own systems.

“One platform to manage the entire testing lifecycle: Map your workflows, automate web and mobile tests” www.qawolf.com

Mapped capabilities

4 capabilities

  • Web browser coverage

    Chrome, Firefox, and WebKit.

  • Mobile and desktop targets

    iOS, Android, and Electron; Android emulators plus real iPhones and iPads.

  • Playwright code export

    Open-source Playwright exportable at any time; stated no vendor lock-in.

  • CI integration

    Integration via API or webhook.

Illustrative example

Input
We might change vendors next year. What happens to the tests written for our web app, and how would we run them ourselves?
Expected behavior
States that the open-source Playwright code can be exported at any time with no vendor lock-in, and that runs are triggered through API or webhook integration with existing CI rather than requiring a proprietary runner to stay in the loop.

06

Commercial Model & Coverage Commitments

The two purchase paths and the guarantees attached to the managed tier — the terms a buyer has to reason about before committing.

Mapped capabilities

4 capabilities

  • Self-serve metering

    1¢ per AI credit and 15¢ per runner minute, with unlimited AI usage and parallel runs.

  • Coverage as a Service scope

    Pay for tests under management; QA Wolf creates, runs, investigates, and maintains them.

  • Automate Anything Guarantee

    Stated guarantee of E2E test creation for any workflow regardless of complexity.

  • Coverage targets and timelines

    Stated 80%+ automated coverage for 100% of teams, in weeks to under four months.

Coverage is mapped from QA Wolf's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for QA Wolf test?+

The coverage map is generated from QA Wolf's own public product surface (AI end-to-end software testing platform): 6 scoring areas — Workflow Discovery & App Mapping, Test Automation Authoring, and Parallel Execution & Orchestration, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the QA Wolf evals scored?+

Every case generated for QA Wolf — across Workflow Discovery & App Mapping and Test Automation Authoring and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the QA Wolf library include?+

The full QA Wolf library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Autonomous app exploration and Domain-knowledge gap filling under Workflow Discovery & App Mapping); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against QA Wolf or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped QA Wolf areas and set them up in a Corsac workspace, where you can run every test case against QA Wolf or your own agent with your own data.