All evals
Momentic

Eval directory

Evals for Momentic

Eval coverage for Momentic, mapped from its public product surface.

About Momentic

Momentic is an AI-powered end-to-end testing platform for web, iOS, and Android. An AI agent turns natural-language descriptions into YAML tests stored in your codebase, which can be run locally, in CI, or in the cloud, with auto-heal, caching, and agents that add coverage and debug failures. A cloud dashboard adds run viewing, analytics, an AI-maintained knowledge base, test quarantining, and enterprise features.

Industry

AI end-to-end testing platform (web and mobile QA)

Use the eval library for Momentic

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Momentic?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Natural-language authoring and test structure

Turning plain-English descriptions into readable YAML tests that live in the codebase, including the step vocabulary, modules, variables, and the declarative-vs-imperative tradeoff.

“Write end-to-end tests for web and mobile in plain English.” momentic.ai

Mapped capabilities

4 capabilities

  • YAML test format, modules, and variables

    Test structure, reuse via modules, variable interpolation such as {{ env.PASSWORD }}.

  • Web and mobile step references

    Selecting the right step for a platform; targets accepting CSS or auto-detected XPath; the fill alias vs. type.

  • AI actions and semantic phases

    Rolling imperative steps into declarative goals, grouping into phases, and designing around stated agent limits.

  • Assertions, including run assertions

    Step-based assertions vs. post-run assertions evaluated against sampled video frames for transient toasts, spinners, and pop-ups.

02

Agentic coverage and exploration

The built-in agent team that adds coverage from a diff, seeds coverage for a whole app, and reports observable product bugs without proposing application code changes.

“Momentic is an AI testing platform that builds, runs, and maintains tests as your product changes.” momentic.ai

Mapped capabilities

4 capabilities

  • Explore from a diff

    momentic ai explore diff over a commit range: identifying user-facing changes, including backend changes that surface to users.

  • Journey mapping and coverage reuse

    Mapping changes to journeys and editing an existing partially-covering test instead of adding a sibling.

  • Build vs. discover modes

    Default live-browser authoring vs. --dry-run discovery; momentic ai explore latest for whole-app seeding.

  • Bug surfacing with evidence

    Severity, ordered reproduction steps, supporting evidence; read-only with respect to application code.

Illustrative example

Input
A backend pricing change landed in abc123..def456. I want the explore agent to list the affected user journeys only — no live browser, no test edits.
Expected behavior
Names momentic ai explore diff abc123..def456 with --dry-run, noting that dry-run discovers journeys without opening a browser session or authoring tests, and that a backend change still counts when it surfaces to users.

03

Execution, caching, and self-healing

How tests actually run and stay green as the product changes: the agent loop, the step cache, auto-heal, and organization-scoped memory that disambiguates natural language.

“Our AI agent turns natural language into reliable, repeatable tests stored as readable YAML files in your codebase.” momentic.ai

Mapped capabilities

4 capabilities

  • Agent loop, auto-heal, and step cache

    Rerun speed from intelligent caching; healing behavior as the app changes.

  • Memory semantics and scope

    Per-organization storage, authenticated runs only, 30-day inactivity expiry, automatic pruning, locator and assertion agents.

  • Memory configuration and caching interaction

    Per-test AI options toggle, ai.useMemory default, and memory being skipped when caching is explicitly disabled.

  • Memory on failed steps

    Writing traces on failed AI-assisted steps so genuine bugs keep failing consistently.

Illustrative example

Input
We turned caching off for this test but left ai.useMemory enabled. Will the locator agent still pull traces from our org's memory on this run?
Expected behavior
Answers no. Memory is treated as a form of caching, so explicitly disabling caching also skips memory, regardless of the ai.useMemory setting; it does not present ai.useMemory as an override.

04

Failure triage, CI gating, and maintenance

What happens after a run goes red: classifying failures in-flight, deciding whether CI blocks, routing failures into AI maintenance, and quarantining unreliable tests.

Mapped capabilities

4 capabilities

  • In-run failure classification

    Classifying failures during a run to separate product bugs from test drift.

  • CI gating control

    Controlling when a failing run blocks a pull request.

  • Routing into AI test maintenance

    Handing failures to the maintenance agents rather than manual repair.

  • Test quarantining

    Dashboard-side quarantine of tests so they stop gating while under repair.

05

Cross-platform setup and environment access

Getting Momentic installed and authenticated across web, iOS, and Android, running the same suite locally, in CI, or in a cloud sandbox, and getting past the app's own login.

Mapped capabilities

4 capabilities

  • Install, wizard, and CLI setup

    npx @momentic/wizard, Node 22.12+/24+, momentic init, browser login vs. MOMENTIC_API_KEY, non-interactive coding-agent path.

  • Platform quickstarts and cross-platform builds

    Chromium web, local/remote iOS simulators, Android emulators for APK builds, native artifacts from React Native, Expo, or Flutter.

  • Application authentication strategies

    Email OTP, SMS OTP, magic links, TOTP, cached sessions, and Vercel preview deployments.

  • Page interactions and test data

    File upload/download, dropdowns, recording and validating network requests, inline fake data, CSV-driven variants.

06

Dashboard, reporting, and enterprise controls

The cloud surface around runs — viewing results, analytics over time, the AI-maintained knowledge base — plus export formats and the account controls an enterprise reviewer checks.

Mapped capabilities

4 capabilities

  • Run results and artifacts

    Runs, videos, traces, local reports, and the editor's execution-following behavior.

  • Analytics and knowledge base

    Trends over time and the AI-maintained knowledge base of terminology and flows the agents draw on.

  • Report export formats

    JSON reports, JUnit output, and self-hosted run results.

  • Plan limits and enterprise governance

    Credit-based run allowances, results retention, mobile device and phone-number limits, SAML SSO, SCIM, audit log, mobile session length.

Coverage is mapped from Momentic's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Momentic test?+

The coverage map is generated from Momentic's own public product surface (AI end-to-end testing platform (web and mobile QA)): 6 scoring areas — Natural-language authoring and test structure, Agentic coverage and exploration, and Execution, caching, and self-healing, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Momentic evals scored?+

Every case generated for Momentic — across Natural-language authoring and test structure and Agentic coverage and exploration and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Momentic library include?+

The full Momentic library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, YAML test format, modules, and variables and Web and mobile step references under Natural-language authoring and test structure); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Momentic or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Momentic areas and set them up in a Corsac workspace, where you can run every test case against Momentic or your own agent with your own data.