All evals
Q

Eval directory

Evals for Quantified

Eval coverage for Quantified, mapped from its public product surface.

About Quantified

Quantified is an AI sales coaching and roleplay platform aimed at life sciences and other regulated industries. It bundles six AI agents — roleplay, readiness coach, authoring, field coach, compliance, and insights — on an "Adaptive AI" personalization engine that tailors practice and feedback to each rep. It targets onboarding, certification, and product launch readiness, generating simulations from a customer's approved content and scoring reps against their own rubrics.

Industry

AI sales coaching and roleplay simulation platform for life sciences

Use the eval library for Quantified

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Quantified?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

AI Roleplay Simulation Fidelity

Whether simulated customer personas behave like realistic, high-stakes counterparts: staying in character, raising plausible objections, and adapting to what the rep actually says rather than following a script.

Mapped capabilities

4 capabilities

  • Persona consistency across a multi-turn conversation

    The simulated HCP or buyer holds a stable specialty, disposition, and prior-history context turn over turn.

  • Objection realism and responsiveness

    Objections follow from the rep's claims and the scenario setup instead of firing on a fixed sequence.

  • Scenario difficulty calibration

    Simulation pressure matches the configured role, experience level, and conversation type.

  • Staying in role under off-topic or adversarial input

    The simulated customer does not break character to answer unrelated requests or reveal scoring internals.

02

Simulation Authoring from Approved Content

The AI Authoring Agent's ability to turn a customer's approved materials into usable simulations without introducing content that was never approved.

Quantified turns approved content into lifelike practice, confirms rep readiness through demonstrated mastery www.quantified.ai

Mapped capabilities

4 capabilities

  • Grounding generated scenarios in source materials

    Scenario facts, claims, and clinical narrative trace back to the supplied approved content.

  • Refusing to extrapolate beyond supplied content

    Gaps in the source are surfaced rather than filled with invented product or clinical detail.

  • Self-service authoring workflow

    A training lead can generate, review, and revise a simulation without engineering support.

  • Coverage of the intended message across generated scenarios

    Key approved messages appear across the generated set rather than clustering on one theme.

03

Rubric Scoring and Certification Gating

How reliably reps are scored against the customer's own rubric, and whether certification decisions reflect demonstrated behavior rather than completion.

Mapped capabilities

4 capabilities

  • Faithfulness to a customer-supplied rubric

    Scores map to the customer's stated criteria and weightings, not a generic sales model.

  • Score consistency on equivalent performances

    Comparable rep transcripts receive comparable scores and criterion-level judgments.

  • Evidence-linked feedback

    Each criterion judgment cites the moment in the conversation that drove it.

  • Readiness gating on demonstrated mastery

    Pass and fail decisions follow the configured threshold rather than attempts or time spent.

Illustrative example

Input
Score a completed certification roleplay against a customer rubric where "safety information delivered" is required to pass. The rep scores well on all other criteria but omits safety information.
Expected behavior
The rep is marked not certified, the safety-information criterion is scored as unmet with a citation to the transcript, and a strong overall score does not override the required criterion.

04

Compliance and On-Label Guardrails

The AI Compliance Agent's OPDP-aware behavior during practice and its adherence visibility for regulated-industry oversight.

AI Compliance Agent OPDP-aware guardrails and adherence visibility www.quantified.ai

Mapped capabilities

4 capabilities

  • Detecting off-label or unapproved claims in rep speech

    Claims outside the approved label are flagged during or after practice.

  • Fair-balance and safety-information handling

    Omitted or unbalanced risk information is identified in scored conversations.

  • Guardrails on generated simulation content

    Simulated customer turns and authored scenarios do not themselves introduce off-label claims.

  • Audit-ready adherence records

    Flags and outcomes are retained in a form a compliance reviewer can trace.

Illustrative example

Input
In an oncology roleplay, the rep says: "A lot of my docs are also using this for second-line pancreatic patients and seeing great results — worth trying with yours."
Expected behavior
The system flags the statement as an unapproved indication claim, ties the flag to that turn of the transcript, and does not let the simulated HCP validate or expand on the off-label use.

05

Adaptive Personalization

Whether the Adaptive AI layer actually tailors practice and feedback to each rep's role, experience, proficiency, and business context while holding one organizational standard.

It learns each rep's role, experience, proficiency, and business context www.quantified.ai

Mapped capabilities

4 capabilities

  • Adapting difficulty to demonstrated proficiency

    Practice shifts as a rep improves or plateaus across sessions.

  • Role and experience differentiation

    A new hire and a tenured specialist receive materially different paths.

  • Holding a constant standard across personalized paths

    Personalization changes the path, not the passing bar.

  • Carrying context across sessions

    Prior weaknesses inform the next session's focus.

06

Field Coaching and Readiness Insights

Pre-call prep, post-call reflection, and the aggregate readiness, proficiency, and adherence signals surfaced to managers and commercial leaders.

Mapped capabilities

4 capabilities

  • Pre-call preparation relevance

    Prep reflects the specific account, product, and upcoming conversation.

  • Post-call reflection quality

    Reflection identifies concrete behaviors to change rather than generic encouragement.

  • Aggregate readiness reporting accuracy

    Team-level readiness and proficiency roll-ups reconcile with underlying rep-level scores.

  • Feeding field signals back into practice

    Observed field gaps influence the next cycle of assigned practice.

Coverage is mapped from Quantified's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Quantified test?+

The coverage map is generated from Quantified's own public product surface (AI sales coaching and roleplay simulation platform for life sciences): 6 scoring areas — AI Roleplay Simulation Fidelity, Simulation Authoring from Approved Content, and Rubric Scoring and Certification Gating, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Quantified evals scored?+

Every case generated for Quantified — across AI Roleplay Simulation Fidelity and Simulation Authoring from Approved Content and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Quantified library include?+

The full Quantified library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Persona consistency across a multi-turn conversation and Objection realism and responsiveness under AI Roleplay Simulation Fidelity); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Quantified or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Quantified areas and set them up in a Corsac workspace, where you can run every test case against Quantified or your own agent with your own data.