All evals
R

Eval directory

Evals for Rehearsals

Eval coverage for Rehearsals, mapped from its public product surface.

About Rehearsals

Rehearsals is a self-serve simulation platform that runs business questions against AI "twins" of real people to predict customer behavior before acting. Twins are built from hour-long voice and video interviews, structured answers, and outcome data, then calibrated to demographic segments. It is positioned as a faster, cheaper alternative to traditional user research for pricing, conversion, ad creative, churn, competitive, and audience questions.

Industry

AI consumer simulation platform for marketing and product research

Use the eval library for Rehearsals

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Rehearsals?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Decision Intake & Study Design

Turning a raw business question into a scoped, runnable Rehearsal: briefing the agent (including by voice), uploading context, defining the decision, and designing the study and simulation set before anything runs.

Each Rehearsal runs thousands of simulations in parallel to predict how people take action or react to change. www.runrehearsals.com

Mapped capabilities

4 capabilities

  • Voice and chat briefing of the agent

    Capturing the business question conversationally and confirming what is being decided.

  • Multimodal context upload

    Accepting videos, decks, product concepts, and rough sketches as study material.

  • Decision framing and study construction

    Translating a vague question into a study the platform designs, including counterfactual/control framings.

  • Simulation scoping and scale

    Selecting which twins and how many parallel simulations a given question warrants.

02

Twin Construction & Segment Calibration

The fidelity claim underneath everything else: twins built inductively from hour-long voice and video interviews, structured answers, and outcome data, then calibrated to real demographic segments and validated against real-world results.

Upload videos, decks, product concepts, or rough sketches to test with twins of real customers. www.runrehearsals.com

Mapped capabilities

4 capabilities

  • Interview-grounded twin depth

    Surfacing the why behind a decision — emotion, context, money, tradeoffs, memory.

  • Demographic calibration to segments

    Composing twin panels that match stated U.S. census segment distributions.

  • Differentiation from synthetic personas

    Connecting outcomes to human reasoning rather than attribute-based assumptions.

  • Continuous validation against outcomes

    Re-testing twin predictions against observed real-world results.

03

Pricing & Willingness to Pay

Mapping what a customer will actually pay as a context-dependent range rather than a number, including how features, framing, and competition move that range.

Mapped capabilities

4 capabilities

  • Willingness-to-pay distribution by segment

    Producing a range per key segment, not a single point estimate.

  • Price-moving features and framings

    Identifying what raises or lowers price tolerance.

  • Price increase pressure-testing

    Simulating reaction to a proposed increase before it ships.

  • Competitive price positioning

    Placing a price against alternatives, grounded in twin behavior.

Illustrative example

Input
At a $2,000 starting price, will a foldable iPhone succeed at launch? Give willingness to pay for our two core segments: existing iPhone upgraders and Android switchers.
Expected behavior
Returns a willingness-to-pay range for each of the two named segments rather than one blended number, and names at least one feature or framing that shifts tolerance up or down, plus where $2,000 sits competitively.

04

Conversion, Creative & Concept Testing

Explaining and pre-testing the path to conversion: why people drop off, whether a new flow or concept will land, and how ad creative and messaging perform with twins of real customers.

Mapped capabilities

4 capabilities

  • Funnel drop-off diagnosis

    Explaining the reason behind a drop-off analytics can only locate.

  • Pre-launch flow validation

    Testing a new or redesigned flow before it goes live.

  • Ad creative evaluation

    Testing creative variants and reactions with twin panels.

  • Concept, messaging, and objection discovery

    Surfacing objections that stop purchase and testing product concepts or offers.

05

Churn, Competitive & Audience Analysis

Decisions about the customer base and the market around it: who is at risk, who is ready to upgrade, which cohorts are being missed, and how to respond when a competitor moves.

Mapped capabilities

4 capabilities

  • Churn risk and win-back drivers

    Why customers leave and what would bring churned customers back.

  • Upgrade and new-tier readiness

    Who is ready to upgrade, and why a subscription tier would or would not land.

  • Competitive move response

    Predicting reaction to a competitor launch and what to do about it.

  • Audience and cohort discovery

    Finding customer cohorts the team is currently missing.

06

Results Transparency & Team Collaboration

How results are inspected, questioned, and acted on: every Rehearsal exposes its simulations, conversations, participant profiles, and reasoning, inside a docs-style surface where the agent stays in the loop.

In less than a day, your team can review results, comment, and ask the agent follow-up questions. www.runrehearsals.com

Mapped capabilities

4 capabilities

  • Full simulation and reasoning trace

    Exposing the simulations, conversations, and participant profiles behind each answer.

  • Docs-style review and commenting

    Team review and comment threads on a returned Rehearsal.

  • Agent-in-the-loop follow-ups

    Tagging the agent in comments to answer questions, spawn simulations, or generate visualizations.

  • Delivery format and turnaround

    Returning results in the clearest actionable format within hours to under a day.

Illustrative example

Input
Checkout conversion fell 12% after our new flow shipped. Which step is losing people and why? Show the simulations behind the answer.
Expected behavior
Identifies the specific step where twins abandon, gives reasons drawn from twin reasoning rather than generic funnel theory, and attaches the supporting simulations, conversations, or participant profiles to each stated reason.

Coverage is mapped from Rehearsals's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Rehearsals test?+

The coverage map is generated from Rehearsals's own public product surface (AI consumer simulation platform for marketing and product research): 6 scoring areas — Decision Intake & Study Design, Twin Construction & Segment Calibration, and Pricing & Willingness to Pay, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Rehearsals evals scored?+

Every case generated for Rehearsals — across Decision Intake & Study Design and Twin Construction & Segment Calibration and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Rehearsals library include?+

The full Rehearsals library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Voice and chat briefing of the agent and Multimodal context upload under Decision Intake & Study Design); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Rehearsals or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Rehearsals areas and set them up in a Corsac workspace, where you can run every test case against Rehearsals or your own agent with your own data.