All evals
TestMu AI

Eval directory

Evals for TestMu AI

Eval coverage for TestMu AI, mapped from its public product surface.

About TestMu AI

TestMu AI, formerly LambdaTest, is an AI-agentic cloud platform for quality engineering that lets teams plan, author, execute, and manage tests across web, mobile, API, database, and performance layers. It bundles products including KaneAI and Kane CLI for test authoring, HyperExecute for scaled remote execution, Test Manager, SmartUI visual testing, Web Scanner, and Agent Testing. It offers a free credit-based tier plus paid and enterprise plans, an MCP server for AI clients, and integrations with tools like Slack, GitHub, and Jira.

Industry

agentic AI software testing cloud

Use the eval library for TestMu AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for TestMu AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agentic Test Authoring (KaneAI & Kane CLI)

Planning and authoring tests from natural language or terminal objectives across UI, API, database, and performance layers, including the auto-heal and vision capabilities bundled with Kane CLI plans.

200 credits per month Local test authoring via CLI Auto-heal & vision included www.testmuai.com

Mapped capabilities

4 capabilities

  • Natural-language test authoring with KaneAI

    Desktop browser tests, mobile app tests, and API testing; supported command types.

  • Kane CLI objectives and generated test cases

    Installation, quick start, writing objectives, generating test cases, CLI reference.

  • Checkpoints and Agent Mode

    Defining checkpoints in a run and when Agent Mode applies.

  • Auto-heal and vision behavior

    What auto-heal and vision cover, and which plans include them.

02

Test Execution Cloud (HyperExecute & Live Testing)

Running authored tests at scale across the agentic cloud, from HyperExecute orchestration to real-time interactive sessions on browsers, real devices, and virtual devices.

The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. www.testmuai.com

Mapped capabilities

4 capabilities

  • HyperExecute configuration and orchestration

    HyperExecute YAML, CLI, GUI (beta), job monitoring, scheduling runs.

  • Framework support for web and app automation

    Selenium, Cypress, Playwright, Puppeteer, K6; Appium, Espresso, XCUI, Flutter.

  • Real-time and cross-platform session testing

    Live web, mobile browser, mobile app, and ChromeOS sessions; browser/OS combination selection.

  • Tunnel, private cloud, and remote execution limits

    Tunnel support, private cloud options, execution minutes tied to plan tier.

03

Test Management & Insights

Organizing testing work in Test Manager and interpreting execution data through dashboards and insight views.

Mapped capabilities

4 capabilities

  • Projects and test case repository

    Creating projects with name, description, tags; importing and creating manual test cases.

  • Test runs, results, and milestones

    Building configured runs, recording results in bulk, tracking milestones and coverage.

  • Migration and Jira issue linking

    1-click migration from other platforms; linking Jira issues to test cases.

  • Insights dashboards

    Pre-built and custom dashboards, widgets, Dashboard CoPilot, flaky test and build insights, command log insights.

04

Visual & Accessibility Scanning (SmartUI & Web Scanner)

Detecting visual regressions and accessibility issues across builds, including how comparisons are configured and how noisy content is handled.

Mapped capabilities

4 capabilities

  • SmartUI screenshot and build comparison

    SDKs, CLI, screenshot upload, build config and comparison options.

  • Advanced comparison and dynamic data handling

    Advanced comparison options, ignoring or normalizing dynamic regions.

  • Smart PDF comparison

    Comparing PDF documents through SmartUI.

  • Web Scanner visual and accessibility scans

    Adding URLs, visual UI scans, accessibility scans, scheduling options.

05

Agent-Facing Interfaces (MCP Server & Agent Testing)

Surfaces built for AI clients and for testing customers' own agents: the remote MCP server exposed to MCP-compatible clients, and the Agent Testing platform.

TestMu AI MCP Server provides five tools, each covering a different area of testing www.testmuai.com

Mapped capabilities

4 capabilities

  • MCP server connection and OAuth setup

    Server URL, client configuration, OAuth flow via testmuai.com.

  • MCP tool selection across products

    HyperExecute, Automation, SmartUI, Accessibility, and Test Manager tools and their scopes.

  • Agent Testing platform coverage

    Supported agent types, testing a first agent, Agent Testing CLI, chat agent API integration.

  • Machine-readable documentation access

    llms.txt index and .md variants of documentation pages for agent consumption.

Illustrative example

Input
My SmartUI run flagged a diff. Which TestMu AI MCP tool debugs that, and how do I authenticate my MCP client?
Expected behavior
Names the SmartUI tool as the one that returns natural-language summaries of pixel, layout, DOM, and perceptual differences, and states that connection uses the remote server URL with an OAuth flow that redirects to testmuai.com on first use.

06

Plans, Credits & Integrations

Credit-based entitlements across free, Starter, Pro, and Enterprise tiers, and the third-party workflow integrations available to freemium and premium accounts.

Slack Integration with TestMu AI is available for freemium as well as premium plan. www.testmuai.com

Mapped capabilities

4 capabilities

  • Credit allocation and reset behavior

    200 free credits per 30 days; 2,000 Starter and 10,000 Pro monthly credits plus launch bonuses.

  • Feature gating by tier

    Test Manager tier, remote HyperExecute execution and minutes, tunnel, scheduling, seats.

  • Enterprise controls

    SSO, advanced access control, IP whitelisting, data retention rules, dedicated support.

  • Workflow integrations setup

    Slack and GitHub integration steps, required admin/user access, pushing bugs and issues; Jira.

Illustrative example

Input
I'm on the free plan. How many credits do I get, and can I run my tests remotely on HyperExecute?
Expected behavior
States the free tier gives 200 credits that reset every 30 days and supports local test authoring via CLI, then says remote execution on HyperExecute is not included at free or Starter and begins with Kane CLI Pro at $99/month with 100 execution minutes.

Coverage is mapped from TestMu AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for TestMu AI test?+

The coverage map is generated from TestMu AI's own public product surface (agentic AI software testing cloud): 6 scoring areas — Agentic Test Authoring (KaneAI & Kane CLI), Test Execution Cloud (HyperExecute & Live Testing), and Test Management & Insights, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the TestMu AI evals scored?+

Every case generated for TestMu AI — across Agentic Test Authoring (KaneAI & Kane CLI) and Test Execution Cloud (HyperExecute & Live Testing) and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the TestMu AI library include?+

The full TestMu AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Natural-language test authoring with KaneAI and Kane CLI objectives and generated test cases under Agentic Test Authoring (KaneAI & Kane CLI)); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against TestMu AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped TestMu AI areas and set them up in a Corsac workspace, where you can run every test case against TestMu AI or your own agent with your own data.