All evals
A

Eval directory

Evals for Autify

Eval coverage for Autify, mapped from its public product surface.

About Autify

Autify is an AI platform for software testing that spans the software testing lifecycle, from test design through execution and maintenance. Its products include Aximo (an autonomous AI testing agent driven by natural language across web, mobile, and desktop), Autify Nexus (AI-powered test automation built on Playwright), Autify Genesis (AI-driven test design from requirements and source code), and an AI-Native Managed QA service. Plans range from a free individual tier to enterprise deployments with on-prem or dedicated infrastructure.

Employees

130 employees globally

Industry

AI software testing / QA automation platform

Website

autify.com

Use the eval library for Autify

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Autify?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Aximo Autonomous Test Execution

The natural-language testing agent that navigates and validates flows directly against a running application without scripts or selectors.

Aximo doesn’t rely on scripts or selectors. It executes directly against your application autify.com

Mapped capabilities

4 capabilities

  • Natural-language scenario interpretation

    Turning a plain-English scenario plus optional start URL into an executed flow with a verified expected outcome.

  • Cross-platform execution

    Web, mobile, and desktop coverage across iOS, Android, Windows, and macOS, including OS-level interactions.

  • Selector-free resilience to UI change

    Executing against the application rather than stored locators, so UI changes do not create maintenance work.

  • Interactive session control

    Starting a session, running with Interactive Mode on, and ending a session mid-run.

Illustrative example

Input
Start URL app.example.com. Test that a registered user can reset their password via the emailed link, then confirm the same link is rejected when used a second time.
Expected behavior
Aximo navigates the reset flow end to end as a real user and reports one clear verdict covering both the successful reset and the rejected reuse, without surfacing scripts or selectors for the tester to maintain.

02

Autify Nexus Test Authoring on Playwright

AI-assisted creation and upkeep of Playwright-based automation, from recording through code export.

Upload your product requirements—PRDs, user stories, or specs—and Nexus translates them into detailed test cases within seconds. autify.com

Mapped capabilities

4 capabilities

  • Natural-language recorder

    Capturing a described user flow as structured, editable test steps.

  • Test case generation from documents

    Translating uploaded PRDs, user stories, and specs into detailed test cases.

  • Fix with AI locator recovery

    Flagging a failed result caused by a changed locator and proposing an alternative.

  • Export to editable Playwright code

    Converting AI-generated or recorded scenarios into full-code Playwright scripts.

Illustrative example

Input
A recorded checkout test fails at the "Place order" step because its locator changed. Request an AI-suggested alternative locator for that step.
Expected behavior
Nexus flags the specific failed step, proposes one alternative locator for it, and asks whether to apply the fix rather than editing silently. Other steps in the scenario remain untouched.

03

Autify Genesis Test Design from Requirements and Code

Enterprise test design that analyzes specifications and codebases to produce features, testing viewpoints, and QA artifacts.

Mapped capabilities

4 capabilities

  • Requirement and source ingestion

    Importing product requirements, source code, documentation, and uploaded files as design inputs.

  • Smart indexing controls

    Honoring .gitignore plus custom patterns and filtering out binaries and oversized files.

  • Feature and testing viewpoint extraction

    Identifying features and viewpoints from analyzed inputs before test case generation.

  • Holistic documentation and in-app chat

    Generating application-spanning documentation from indexed code and answering questions against it.

04

Results, Maintenance, and Continuous Testing

What a run reports afterward and how that output feeds reuse, maintenance, and CI/CD.

Mapped capabilities

4 capabilities

  • Verdict and run summary

    Reporting whether the scenario behaved as expected, with start, completion, and duration.

  • Session traceability views

    Summary, Logs, Console, and Conversation records for an executed session.

  • Save as case and rerun

    Promoting a session into a reusable case and re-executing it.

  • Automation code generation for CI/CD

    Producing automation code from designed tests and running it in a CI/CD pipeline.

05

Plans, Credits, and Concurrency

The metering and entitlement surface that determines what a given account can run and how much it consumes.

Mapped capabilities

4 capabilities

  • Credit consumption and model multiplier

    Per-session credit spend tied to the selected model, against remaining balance.

  • Tiered feature gating

    Capabilities such as Script Generation and GenAI available only at qualifying plan tiers.

  • Web and mobile concurrency limits

    Enforcing per-plan parallel session ceilings separately for web and mobile.

  • Billing terms and add-on credits

    Annual versus monthly pricing, allotted credits, and purchasing additional credits.

06

Deployment, Access, and Data Boundaries

Where the platform runs and what it is permitted to touch, which drives enterprise procurement.

Mapped capabilities

4 capabilities

  • Hosting model

    Cloud hosting for lower tiers versus on-prem or dedicated infrastructure for enterprise.

  • Network access controls

    IP whitelisting for reaching protected environments under test.

  • Local scope and secret storage

    Desktop operation limited to user-selected folders with secrets held in Keychain.

  • Regional site routing

    Offering visitors from Korea a redirect to Autify NoCode while preserving the corporate-site choice.

Coverage is mapped from Autify's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Autify test?+

The coverage map is generated from Autify's own public product surface (AI software testing / QA automation platform): 6 scoring areas — Aximo Autonomous Test Execution, Autify Nexus Test Authoring on Playwright, and Autify Genesis Test Design from Requirements and Code, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Autify evals scored?+

Every case generated for Autify — across Aximo Autonomous Test Execution and Autify Nexus Test Authoring on Playwright and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Autify library include?+

The full Autify library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Natural-language scenario interpretation and Cross-platform execution under Aximo Autonomous Test Execution); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Autify or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Autify areas and set them up in a Corsac workspace, where you can run every test case against Autify or your own agent with your own data.