All evals
F

Eval directory

Evals for Functionize

Eval coverage for Functionize, mapped from its public product surface.

About Functionize

Functionize Studio is an agentic quality layer that independently tests AI-written code against a customer's running application. It writes tests from plain-language prompts, runs them across browsers and environments in parallel, self-heals tests when the UI changes, and turns runs into a release signal while the customer sets the quality standard. It is sold self-serve from a free tier through individual, team, and custom enterprise plans, alongside Functionize's broader enterprise automation platform.

Industry

agentic AI software test automation (QA) platform

Use the eval library for Functionize

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Functionize?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Prompt-to-Test Authoring

Turning plain-language descriptions of a feature or flow into runnable tests that join the existing suite, with mapped test data and minimal instruction as flows grow more complex.

Machine learning-based tests use big data to understand website updates and self-heal workflows www.functionize.com

Mapped capabilities

4 capabilities

  • Plain-language test generation

    Builds a test from a described feature or user flow without scripted steps.

  • Test data mapping

    Associates the data a generated flow needs so the test can execute end to end.

  • Suite integration

    Adds new coverage to the existing suite rather than producing a standalone artifact.

  • Prompt economy on complex flows

    Handles longer flows without proportionally more instruction from the user.

02

Execution Across Browsers and Environments

Running the suite in parallel across browsers, environments, and geographies, with concurrency bounded by plan, and completing in minutes rather than hours.

Mapped capabilities

4 capabilities

  • Parallel run scheduling

    Distributes runs concurrently up to the plan's parallel-run ceiling.

  • Cross-browser coverage

    Executes the same flow across supported browsers.

  • Environment and geolocation targeting

    Runs against a chosen environment or geographic context.

  • Live debug and run inspection

    Exposes an in-flight or replayable view of what a run did.

03

Self-Healing and UI Change Resilience

Re-identifying elements when the UI moves, using a multi-hundred-datapoint element model, so a changed interface produces a report instead of a repair backlog.

Every run trains Studio on your app, so coverage sharpens with use. www.functionize.com

Mapped capabilities

4 capabilities

  • Element re-identification

    Resolves a moved or restyled element from its collected attributes.

  • Mid-run healing

    Repairs the affected test during the run rather than failing it outright.

  • Healing transparency

    Reports what was healed and why so the change is reviewable.

  • Coverage that sharpens with use

    Uses prior runs on the same app to improve subsequent identification.

Illustrative example

Input
Our signup test broke after last night's deploy — the Continue button moved and got renamed. What happened, and do I need to fix the test?
Expected behavior
Explains that Studio re-identified the element and healed the affected test mid-run, flags it as a UI change rather than a product defect, and points to the run report. Does not ask the user to hand-update selectors.

04

Failure Diagnosis and Release Signal

Separating genuine defects from noise, diagnosing why a run failed, and converting the result into a clear pass/fail signal against a standard the customer sets.

Mapped capabilities

4 capabilities

  • Real bug vs. noise triage

    Classifies a failure as a product defect or an environmental/UI artifact.

  • Failure diagnosis

    Explains the cause of a failed step in reviewable terms.

  • Release gating on a user-set standard

    Applies the customer's definition of ready; the release call stays with the user.

  • Run-level reporting

    Summarizes a run into evidence a team can act on.

05

Test Data, Channels, and Programmatic Access

Covering flows that leave the browser — email and SMS verification, MFA/OTP, activation links — and driving Studio from outside the UI where the plan permits.

Studio resolves each element from 200+ data points, with models that carry no other job www.functionize.com

Mapped capabilities

4 capabilities

  • Email testing for MFA and activation links

    Completes flows requiring an emailed code or link.

  • SMS testing for MFA/OTP

    Completes flows requiring a texted one-time code.

  • API and MCP access

    Invokes Studio programmatically on plans that include it.

  • API Explorer

    Interactive surface for exercising the API.

06

Plans, Credits, and Governance

Self-serve tiers from free through team and enterprise, with credit budgets, parallelism ceilings, retention windows, workspace and access controls, and billing arrangements that differ by plan.

Mapped capabilities

4 capabilities

  • Credit budgets and consumption

    Monthly included credits per plan, including pooled enterprise credits.

  • Retention and quota limits

    Data retention window and parallel-run cap that apply to the current plan.

  • Workspace, SSO, and RBAC

    Single vs. multi-team workspaces and enterprise access controls.

  • Audit logs and usage analytics

    Administrative visibility into activity and consumption.

Illustrative example

Input
I'm on the Free plan. Can I run eight browser checks at once and keep three months of run history?
Expected behavior
States that Free allows at most 5 parallel runs and 1 month of data retention, so neither request fits, and directs the user to a higher plan for more concurrency and longer retention.

Coverage is mapped from Functionize's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Functionize test?+

The coverage map is generated from Functionize's own public product surface (agentic AI software test automation (QA) platform): 6 scoring areas — Prompt-to-Test Authoring, Execution Across Browsers and Environments, and Self-Healing and UI Change Resilience, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Functionize evals scored?+

Every case generated for Functionize — across Prompt-to-Test Authoring and Execution Across Browsers and Environments and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Functionize library include?+

The full Functionize library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Plain-language test generation and Test data mapping under Prompt-to-Test Authoring); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Functionize or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Functionize areas and set them up in a Corsac workspace, where you can run every test case against Functionize or your own agent with your own data.