All evals
Amplitude

Eval directory

Evals for Amplitude

Eval coverage for Amplitude, mapped from its public product surface.

About Amplitude

Amplitude is a product analytics platform that combines product analytics, session replay, experimentation, guides and surveys, and AI feedback in one interface. It offers a conversational "Global Agent" plus an MCP server and SDKs so AI clients and developer tools can query analytics, taxonomy, and content. Plans range from a free tier with 2M events/month to custom enterprise event-based pricing, with unlimited seats on every plan.

Industry

AI-powered product analytics platform

Use the eval library for Amplitude

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Amplitude?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Global Agent and Conversational Analytics

The natural-language agent that answers product questions by pulling from what Amplitude already knows, including root-cause explanations and follow-up refinement.

Amplitude is the highest performing analytics agent on the market amplitude.com

Mapped capabilities

4 capabilities

  • Question-to-analysis translation

    Turning an open product question into the correct analysis type (funnel, retention, segmentation) and time window.

  • Grounded root-cause explanation

    Attributing a drop-off or change to a specific step or segment with the counts behind it.

  • Multi-turn follow-up and refinement

    Preserving prior scope when the user narrows, re-segments, or asks a suggested follow-up.

  • Scope and unavailability handling

    Saying what cannot be answered from instrumented data instead of producing an unsupported number.

Illustrative example

Input
Why are new users dropping off before they finish setup this week?
Expected behavior
Identifies the signup-to-setup funnel, states the drop-off with the underlying counts and the time window used, and attributes it to a specific step or segment rather than a generic guess.

02

MCP Server and Agent Tooling

Amplitude MCP exposing analytics, taxonomy, and content as tools callable by AI clients, plus the setup and access-scoping around it.

Amplitude MCP exposes Amplitude analytics, taxonomy, and content as tools that an AI client amplitude.com

Mapped capabilities

4 capabilities

  • Progressive tool discovery

    Discovering the right tool before calling it rather than assuming tool names or arguments.

  • Analysis workflow over MCP

    Querying charts, dashboards, cohorts, metrics, and Session Replay data through tool calls.

  • Creation and editing workflow

    Creating or editing dashboards, notebooks, experiments, flags, guides, and surveys via tools.

  • Client setup and access scope

    Connecting supported clients, selecting the correct region, and respecting role-scoped MCP access.

Illustrative example

Input
Using the Amplitude MCP server, find events in our tracking plan that fired fewer than 100 times last month, then build a dashboard from them.
Expected behavior
Discovers the taxonomy and dashboard tools before calling them instead of assuming tool names, reads the tracking plan first, creates the dashboard from only those events, and reports the created dashboard identifier.

03

Core Product Analytics

The self-serve analytics workhorse: funnels, retention, cohorts, and the reporting objects teams build on top of them.

Turn your questions into insights faster with self-serve product analytics. amplitude.com

Mapped capabilities

4 capabilities

  • Funnel and drop-off analysis

    Building conversion funnels and locating the step where users are lost.

  • Retention and behavioral cohorts

    Defining behavior-based cohorts and reading retention over them.

  • Dashboards, notebooks, and templates

    Assembling and sharing reporting surfaces, including out-of-the-box metric templates.

  • Custom events, formulas, and alerts

    Deriving computed events and metrics and configuring monitoring on them.

04

Session Replay and Experience Analytics

Qualitative layer that connects a number in a chart to the actual sessions, clicks, and page regions behind it.

Auto-instrument page views, clicks, and sessions out of the box. amplitude.com

Mapped capabilities

4 capabilities

  • Finding friction sessions

    Locating replays matching a described stuck or error behavior.

  • Quantitative-to-qualitative linking

    Moving from a funnel step or cohort to the replays of those users.

  • Heatmaps and zoning insights

    Reading click, scroll, and page-zone engagement on a given surface.

  • Replay capture limits and privacy

    Behavior at monthly replay quotas and around sensitive captured content.

05

Experimentation, Guides, and Feedback

Acting on insight: feature and web experiments, flag-driven rollout, in-product guides and surveys, and the AI feedback loop.

Mapped capabilities

4 capabilities

  • Experiment design and readout

    Setting up a feature or web experiment and interpreting its result correctly.

  • Feature flag targeting and rollout

    Targeting flags to cohorts and staging exposure.

  • Guides and surveys delivery

    Configuring in-product guides and surveys and their audience conditions.

  • AI feedback synthesis

    Turning collected feedback records into opportunities tied to product behavior.

06

Instrumentation, Taxonomy, and Plan Boundaries

Getting data in correctly and knowing what the account is entitled to: SDKs across platforms, the shared identity and event model, tracking plan hygiene, and plan limits.

Every SDK shares the same identity, event, and consent model amplitude.com

Mapped capabilities

4 capabilities

  • SDK install and initialization

    Correct install and init guidance for the requested platform SDK and product (Analytics, Experiment, Replay, Guides).

  • Identity, event, and consent model

    Consistent user identification, user/group properties, and consent handling across SDKs.

  • Tracking plan and taxonomy changes

    Inspecting and amending the tracking plan without breaking existing charts.

  • Plan tiers and entitlements

    Accurate statements about event volumes, replay and feedback quotas, seats, and admin features per plan.

Coverage is mapped from Amplitude's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Amplitude test?+

The coverage map is generated from Amplitude's own public product surface (AI-powered product analytics platform): 6 scoring areas — Global Agent and Conversational Analytics, MCP Server and Agent Tooling, and Core Product Analytics, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Amplitude evals scored?+

Every case generated for Amplitude — across Global Agent and Conversational Analytics and MCP Server and Agent Tooling and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Amplitude library include?+

The full Amplitude library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Question-to-analysis translation and Grounded root-cause explanation under Global Agent and Conversational Analytics); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Amplitude or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Amplitude areas and set them up in a Corsac workspace, where you can run every test case against Amplitude or your own agent with your own data.