All evals
T

Eval directory

Evals for testRigor

Eval coverage for testRigor, mapped from its public product surface.

About testRigor

testRigor is a generative AI test automation tool that lets teams write end-to-end tests in plain English instead of code, so tests don't break when UI locators change. It covers web, mobile (native and hybrid iOS/Android), native desktop, mainframe, API, email, SMS/phone call, and 2FA testing. Manual test cases can be copy-pasted or imported and then corrected or expanded using built-in plain English commands.

Industry

AI-based test automation (codeless QA)

Use the eval library for testRigor

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for testRigor?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Plain English Test Authoring

Interpreting free-flowing English instructions and turning them into executable steps, including importing and correcting existing manual test cases.

With testRigor, you can use free-flowing plain English to build test automation. testrigor.com

Mapped capabilities

4 capabilities

  • High-level intent decomposition

    Expanding an instruction like 'purchase a Kindle' into ordered concrete steps (search, enter, click).

  • Manual test case import and correction

    Copy-pasted or imported manual cases refined with supported built-in commands.

  • Built-in command coverage

    Basic commands, comments, referencing locations, and element/screen structure interactions.

  • Reusable rules (subroutines)

    Custom and built-in reusable rules, including the Salesforce-specific built-ins.

Illustrative example

Input
Author a single test step reading: purchase a Kindle. Run it against a storefront with a search box, results list, and cart.
Expected behavior
The instruction is decomposed into concrete ordered actions — enter the search term, submit, select the product, add to cart — and each generated step is shown to the author for correction rather than executed opaquely.

02

Cross-Platform Execution Surfaces

Running the same plain-English tests against the platforms testRigor documents support for.

Cover both cross-browser and cross-platform scenarios within a single test. testrigor.com

Mapped capabilities

4 capabilities

  • Web and mobile web

    Cross-browser and cross-platform scenarios in a single test on Windows, macOS, Ubuntu, iOS, and Android.

  • Native and hybrid mobile apps

    iOS and Android app testing, including LambdaTest/BrowserStack device expansion.

  • Native desktop applications

    Windows desktop testing, documented as paid-tier only.

  • Mainframe applications

    Mainframe test creation and execution via the same authoring language.

03

Non-UI and Multi-Channel Testing

Verifying flows that leave the browser: APIs, email, telephony, and authentication challenges.

Mapped capabilities

4 capabilities

  • API testing

    Invoking APIs, retrieving values, validating return codes, storing saved values.

  • Email testing

    Sending email with attachments and verifying deliverability and content.

  • SMS and phone calls

    Twilio integration for placing/verifying calls, sending SMS, confirming deliverability.

  • 2FA login flows

    Completing two-factor authentication logins, including SMS-delivered codes.

Illustrative example

Input
Log in with valid credentials to an account whose second factor is an SMS one-time code sent to a Twilio-backed number.
Expected behavior
The test submits credentials, waits for the SMS to arrive, extracts the one-time code from the message body, enters it in the challenge field, and reaches the authenticated landing page.

04

Validation and Data Handling

Assertions and the data plumbing that supports them across visual, tabular, and generated data.

Mapped capabilities

4 capabilities

  • Validations and Boolean logic

    Validation commands combined with AND/OR/NOT and bracketed expressions.

  • Visual and AI-based checks

    Visual testing and AI-based testing commands, plus audio testing.

  • Table and query operations

    Selecting rows, calculating aggregates, and database query steps.

  • Data extraction and generation

    Variables, loops, conditionals, JsonPath, regex, simple templates, and unique data generation.

05

Test Stability and Maintenance

The product's central claim that tests survive UI change, reducing maintenance effort over time.

Build test automation 100X faster and spend 200X less time maintaining it! testrigor.com

Mapped capabilities

4 capabilities

  • Locator-independent resilience

    Tests written from the end user's perspective continuing to pass after locator changes.

  • AI-based self-healing

    Documented self-healing behavior, including the Selenium self-healing positioning.

  • Ambiguity and failure reporting

    Behavior when an instruction cannot be resolved to an on-screen element or step.

  • Environment-state controls

    Cookies, localStorage, sessionStorage, and userAgent manipulation for reproducible runs.

06

Integrations and Developer Workflow

Getting tests created, run, and maintained from outside the web UI.

Mapped capabilities

4 capabilities

  • Command line execution

    Documented CLI entrypoint for running tests.

  • MCP server and Claude Code

    Writing, running, and maintaining testRigor tests through the MCP server integration.

  • Test recorder and file uploads

    Recorder-assisted authoring plus file upload steps.

  • Specialized environment handling

    Chrome extension testing, captcha resolution, QR code scanning, and advanced JavaScript.

Coverage is mapped from testRigor's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for testRigor test?+

The coverage map is generated from testRigor's own public product surface (AI-based test automation (codeless QA)): 6 scoring areas — Plain English Test Authoring, Cross-Platform Execution Surfaces, and Non-UI and Multi-Channel Testing, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the testRigor evals scored?+

Every case generated for testRigor — across Plain English Test Authoring and Cross-Platform Execution Surfaces and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the testRigor library include?+

The full testRigor library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, High-level intent decomposition and Manual test case import and correction under Plain English Test Authoring); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against testRigor or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped testRigor areas and set them up in a Corsac workspace, where you can run every test case against testRigor or your own agent with your own data.