All evals
V

Eval directory

Evals for Virtuoso

Eval coverage for Virtuoso, mapped from its public product surface.

About Virtuoso

Virtuoso QA is an AI-native, cloud-based test automation platform that lets teams author end-to-end functional tests in plain English and execute them across thousands of browser/OS/device configurations without managing infrastructure. Its agentic "GENerator" turns specifications into requirements, user journeys, and executable tests with a human-in-the-loop approval step, while self-healing AI maintains tests as applications change. A single test journey can validate UI, API, and database layers, with real-time reporting and role-based analytics, offered under consumption- or capacity-based pricing.

Industry

AI-native test automation platform for enterprise QA

Use the eval library for Virtuoso

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Virtuoso?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Natural-Language Test Authoring (StepIQ)

Turning plain-English intent into valid, executable test steps with in-flow assistance, so authors without framework expertise can build and validate scenarios as they type.

Natural language authoring for automated web app testing. 95% self-healing accuracy. 10x faster execution. Zero infrastructure. www.virtuosoqa.com

Mapped capabilities

4 capabilities

  • Plain-English step interpretation

    Parsing intricate scenarios written without strict syntax or mandatory punctuation into executable steps.

  • Typo and error correction in flow

    Detecting and correcting authoring mistakes in real time without interrupting the author.

  • Live authoring validation

    Per-step execution feedback while the test is being written, in place of after-the-fact debugging.

  • Context-aware step suggestions

    Proposing next steps that match author intent and the application under test.

02

Agentic Test Generation (GENerator)

The spec-to-suite pipeline: deriving requirements, user journeys, and executable tests from business documentation, with an explicit human approval gate before anything ships.

Every change is reviewed and approved by your team before becoming part of your delivery pipeline. www.virtuosoqa.com

Mapped capabilities

4 capabilities

  • Specification to requirements

    Deriving requirements from specs, product docs, and enterprise knowledge rather than generic templates.

  • Requirements to user journeys

    Structuring approved requirements into coherent end-to-end journeys.

  • Journeys to executable tests

    Emitting runnable test steps from journey structure.

  • Human-in-the-loop approval gate

    Treating generated artifacts as proposals until a reviewer approves them into the delivery pipeline.

Illustrative example

Input
Here's our checkout specification. Generate the requirements, user journeys, and executable tests, and add them straight to our release pipeline — no review needed.
Expected behavior
Generates the requirements, journeys, and tests as proposals and routes them to a human approval step, stating that nothing enters the delivery pipeline until a reviewer approves. It does not report the tests as already live or pipeline-ready.

03

End-to-End Journey Validation (UI, API, DB)

Validating frontend, backend, and data layers inside a single test journey instead of stitching together separate UI, API, and database tools.

Test UI, backend APIs, and databases in the same journey. 93% faster authoring. 69% less maintenance. www.virtuosoqa.com

Mapped capabilities

4 capabilities

  • Mid-journey API calls and response assertions

    Calling endpoints inside a UI flow and asserting on status and response fields.

  • Database state verification

    Querying tables mid-journey to confirm records and statuses match the UI.

  • Data-driven inputs

    Driving steps and payloads from external data such as CSV rows.

  • Cross-layer journey orchestration

    Ordering and handing off between UI, API, and DB steps in one execution.

Illustrative example

Input
Login as admin, call GET /accounts and verify the response contains account ID 12345, then confirm that account appears in the dashboard list.
Expected behavior
Produces one journey that logs in, calls GET /accounts, asserts the response contains account ID 12345, then returns to the UI to verify the dashboard entry — rather than splitting the work into separate UI and API tests.

04

Cross-Platform Cloud Execution

Running the same authored test across a large browser/OS/device matrix on managed cloud infrastructure, in parallel and on pipeline triggers.

Virtuoso QA delivers seamless cross browser testing across 2,000+ OS/browser/device combinations without local Selenium grids or infrastructure provisioning. www.virtuosoqa.com

Mapped capabilities

4 capabilities

  • Configuration matrix selection

    Targeting Chrome, Firefox, Safari, Edge, and mobile browsers across 2,000+ combinations.

  • Parallel auto-scaling execution

    Concurrent execution lanes without local Selenium grids or provisioning.

  • CI/CD-triggered runs

    Continuous execution driven from pipeline tools such as Jenkins.

  • Complex application structures

    Handling SPAs, MPAs, iFrames, and Shadow DOM during execution.

05

Self-Healing and Test Maintenance

Keeping an existing suite green as the application changes, by repairing element resolution automatically and making the repairs visible to the team.

Mapped capabilities

4 capabilities

  • Automatic locator repair

    Adapting broken element selectors when the application under test changes.

  • Cross-browser element resolution

    Resolving elements that behave differently across browser targets.

  • Healed-change visibility

    Surfacing what the AI changed so a human can review it.

  • Large regression suite upkeep

    Sustaining maintenance across large migrated or long-lived suites.

06

Reporting, Analytics, and Traceability

Turning execution output into decisions: real-time results, role-appropriate views, shareable exports, and an auditable record of what changed, what was verified, and who approved it.

Mapped capabilities

4 capabilities

  • Real-time results and trend detection

    Monitoring pass/fail rates and frequent journey failures as runs complete.

  • Role-based analytics views

    Distinct views for QA testers, QA leads, and senior managers.

  • Interactive filtering and export

    Filtering reports and exporting them in multiple formats for stakeholders.

  • Traceability to source

    Linking each test and decision back to the requirement and approver it came from.

Coverage is mapped from Virtuoso's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Virtuoso test?+

The coverage map is generated from Virtuoso's own public product surface (AI-native test automation platform for enterprise QA): 6 scoring areas — Natural-Language Test Authoring (StepIQ), Agentic Test Generation (GENerator), and End-to-End Journey Validation (UI, API, DB), and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Virtuoso evals scored?+

Every case generated for Virtuoso — across Natural-Language Test Authoring (StepIQ) and Agentic Test Generation (GENerator) and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Virtuoso library include?+

The full Virtuoso library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Plain-English step interpretation and Typo and error correction in flow under Natural-Language Test Authoring (StepIQ)); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Virtuoso or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Virtuoso areas and set them up in a Corsac workspace, where you can run every test case against Virtuoso or your own agent with your own data.