All evals
M

Eval directory

Evals for mabl

Eval coverage for mabl, mapped from its public product surface.

About mabl

mabl is an AI-native testing platform that autonomously builds, runs, maintains, and analyzes end-to-end tests across web, mobile, and API surfaces. It pulls context from code bases to generate tests, executes them in parallel from commit through staging, and uses GenAI to auto-heal tests when the application changes. Failure triage and quality data are pushed into developer workflows such as Jira, the mabl CLI, and coding agents like Claude Code.

Industry

agentic AI software test automation platform

Use the eval library for mabl

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for mabl?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Autonomous Test Creation

Generating new end-to-end coverage by pulling context from the application and its code base, rather than having humans hand-author each test.

mabl independently verifies your application with the proven harness, application context, and agentic intelligence www.mabl.com

Mapped capabilities

4 capabilities

  • Codebase-context test generation

    Proposing tests derived from application and repository context

  • End-to-end user journey coverage

    Composing multi-step flows such as login through checkout

  • Data-driven test variants

    Parameterizing a flow across input sets

  • Reusable flow composition

    Factoring shared steps into flows referenced by multiple tests

02

Execution & CI/CD Orchestration

Running the suite on demand across the delivery pipeline, from early commits through staging deployments, with cloud and local execution paths.

Unlimited apps, environments, and workspaces Unlimited test run concurrency Unlimited local and CI test runs www.mabl.com

Mapped capabilities

4 capabilities

  • Parallel and concurrent runs

    Fanning a suite out across concurrent cloud executions

  • Pipeline-triggered execution

    Invoking runs at commit, PR, and staging deployment points

  • Local and CLI-driven runs

    Executing from the mabl CLI and developer machines

  • Environment and workspace targeting

    Directing runs at the correct app, environment, and workspace

03

Auto-Healing & Test Maintenance

Keeping existing coverage current when the application under test changes, using GenAI to repair tests in real time instead of requiring manual rework.

When apps change, mabl uses GenAI to solve for complex fixes in real time, updating tests without manual rework. www.mabl.com

Mapped capabilities

4 capabilities

  • Selector and locator repair

    Recovering steps when UI elements move or change

  • Agentic runtime recovery

    Recovering an in-flight run rather than failing outright

  • Heal-versus-fail discrimination

    Declining to heal when the change is a genuine regression

  • Change transparency

    Surfacing what was auto-updated and why

Illustrative example

Input
A checkout test clicks "Place Order." In one build the button's label and DOM id changed; in another build the button was removed entirely from the page.
Expected behavior
The renamed button is auto-healed and the run continues, with the update reported. The removed button is not healed: the run fails and is reported as a genuine application regression rather than a test defect.

04

Failure Triage & Developer Handoff

Analyzing every failure, attributing root cause, and pushing the resulting insight into the tools developers already work in.

Failures are triaged by mabl with insights pushed to Jira, your CLI, or tools like Claude Code. www.mabl.com

Mapped capabilities

4 capabilities

  • Root cause insight

    Attributing a failure to a probable underlying cause

  • Automatic failure summaries

    Producing a readable summary of what broke

  • Workflow delivery to Jira, CLI, and coding agents

    Routing triage output into developer destinations

  • Diagnostics available to external agents

    Exposing run diagnostics for consumption outside mabl

Illustrative example

Input
A nightly regression run fails on three tests that all traverse the same login step after an auth service deploy.
Expected behavior
The failures are triaged to a shared root cause at the login step rather than reported as three unrelated defects, and the summary with supporting diagnostics is pushed to the configured Jira project.

05

Full-Stack Coverage Breadth

Validating the whole application surface — UI and API, web and mobile — plus adjacent checks the platform advertises as included capabilities.

Mapped capabilities

4 capabilities

  • Web and cross-browser UI

    Exercising browser-based user workflows

  • Mobile app testing

    Covering iOS and Android from the same platform

  • API contract and breaking-change detection

    Catching API breaks before the frontend is affected

  • Accessibility, performance, visual, email, and document checks

    Non-functional and content assertions bundled into a run

06

Quality Data & Enterprise Governance

Consolidating quality signal across surfaces for leadership visibility, under the access, audit, and support posture enterprises require.

Mapped capabilities

4 capabilities

  • Cross-surface quality dashboards

    Unifying web, mobile, and API results into one view

  • Audit trails

    Recording who changed or ran what

  • Workspace and access scoping

    Separating apps, environments, and workspaces across teams

  • Integration fan-out to Slack, Teams, and Jira

    Broadcasting quality state into team channels

Coverage is mapped from mabl's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for mabl test?+

The coverage map is generated from mabl's own public product surface (agentic AI software test automation platform): 6 scoring areas — Autonomous Test Creation, Execution & CI/CD Orchestration, and Auto-Healing & Test Maintenance, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the mabl evals scored?+

Every case generated for mabl — across Autonomous Test Creation and Execution & CI/CD Orchestration and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the mabl library include?+

The full mabl library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Codebase-context test generation and End-to-end user journey coverage under Autonomous Test Creation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against mabl or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped mabl areas and set them up in a Corsac workspace, where you can run every test case against mabl or your own agent with your own data.