All evals
N

Eval directory

Evals for Nomi

Eval coverage for Nomi, mapped from its public product surface.

About Nomi

Nomi is a real-time AI sales copilot that overlays live video calls and suggests the next best response as the conversation happens. It surfaces competitive battle cards, diagnoses prospect signals, and builds playbooks learned from top-performing reps. It is sold per-seat across Professional, Enterprise, and custom Platform tiers, with CRM and Zoom integrations.

Industry

real-time AI sales copilot

Use the eval library for Nomi

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Nomi?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Live Call Guidance

The core in-call loop: detect a signal in the prospect's speech, diagnose the underlying concern, and propose next-best responses with a stated rationale, fast enough to be usable mid-conversation.

When a prospect mentions a competitor, Nomi surfaces the best angle to take www.nomi.so

Mapped capabilities

4 capabilities

  • Signal detection from live speech

    Identifying the salient moment (e.g. timeline risk, budget hesitancy) from an in-progress utterance rather than the whole call.

  • Diagnosis quality

    Naming the core concern behind a prospect statement instead of restating the surface words.

  • Response option generation

    Offering usable phrasings plus a 'why' that ties the suggestion to a concrete next step.

  • Latency and turn discipline

    Guidance arriving within the tier's stated latency budget and respecting the per-turn suggestion count.

02

Competitive Battle Cards

Behavior when a prospect names a competitor: surfacing the recommended angle, the supporting evidence, and a confidence signal without inventing comparative claims.

Mapped capabilities

4 capabilities

  • Competitor mention triggering

    Recognizing named rivals (Salesforce, Gong, Chorus, Otter, Fireflies) and pulling the matching card.

  • Angle selection and confidence

    Choosing a winning angle and reporting confidence rather than asserting certainty.

  • Evidence grounding

    Backing angles only with claims present in the card; no fabricated benchmarks, pricing, or customer counts.

  • Handling unknown competitors

    Graceful behavior when no card exists for the named tool.

Illustrative example

Input
Live call, prospect says: "We like the platform, but we're also evaluating Salesforce — how does your implementation timeline compare to theirs?"
Expected behavior
Surfaces the Salesforce battle card with a speed-to-value angle and the evidence backing it, and offers a question that moves toward a concrete next step. Does not assert specific Salesforce implementation timelines, pricing, or benchmark figures that are not on the card.

03

Playbooks and Coaching Rollout

Learning patterns from top-performing reps' conversations and turning them into playbook steps that can be reviewed and deployed to a team.

Mapped capabilities

4 capabilities

  • Pattern extraction from call corpora

    Deriving a stated pattern and its win-rate correlation from a set of conversations.

  • Playbook step generation

    Turning a detected pattern into a concrete stage-scoped step (e.g. discovery-phase actions).

  • Attribution and provenance

    Marking machine-added steps and the evidence base behind them.

  • Team deployment scope

    Applying playbook changes to the intended team without silently overwriting manual edits.

04

Plans and Entitlements

Correctly representing and enforcing the Professional, Enterprise, and Platform tiers — including gated features, suggestion limits, latency targets, analytics depth, and support and SLA commitments.

Mapped capabilities

4 capabilities

  • Feature gating by tier

    Battle cards, playbook builder, deal risk alerts, custom model training, and agent personalization gated to the correct plan.

  • Pricing and billing accuracy

    Per-seat prices, annual-vs-monthly framing, trial, and cancellation answers matching published terms.

  • Upgrade path messaging

    Naming the tier that unlocks a requested capability instead of refusing flatly or overpromising.

  • Plan-name disambiguation

    Resolving mismatches between published tier names and names used in FAQ or sales conversation.

Illustrative example

Input
A rep on the Professional plan asks Nomi in-app: "Can you build me a playbook from my team's last quarter of calls and set up deal risk alerts?"
Expected behavior
Declines to perform the action on the current plan, states that playbook builder and deal risk alerts are Enterprise features, and points to the upgrade path. Does not partially execute the request or imply the capability is available on Professional.

05

Integrations and Meeting Setup

Connecting Nomi to conferencing and CRM systems and guiding users through installation, permissions, and supported-platform questions.

Nomi floats on top of Zoom, Meet, or Teams and helps you diagnose what's happening www.nomi.so

Mapped capabilities

4 capabilities

  • Zoom app installation and overlay setup

    Walking a user through connecting Nomi to Zoom and clarifying what must be installed.

  • CRM connection (Salesforce, HubSpot, Attio)

    Connecting an account and explaining what data flows in each direction.

  • Platform coverage answers

    Accurate statements about Meet and Teams support, calendar integration, and language coverage.

  • In-call CRM capture prompts

    Identifying CRM-critical information during a call and prompting the rep to capture it.

06

Security, Data, and Identity

Handling of call recordings and transcripts, enterprise access controls, deployment options, and disambiguation of Nomi.so from the unrelated Nomi.ai product.

Mapped capabilities

4 capabilities

  • Recording and transcript handling

    Explaining what is captured, stored, and retained during live calls.

  • Enterprise access controls

    SSO, audit logs, and on-premise deployment availability by tier.

  • Compliance claim discipline

    Stating only compliance and SLA commitments that appear in published material.

  • Brand and identity disambiguation

    Distinguishing Nomi.so, the sales copilot, from the similarly named Nomi.ai.

Coverage is mapped from Nomi's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Nomi test?+

The coverage map is generated from Nomi's own public product surface (real-time AI sales copilot): 6 scoring areas — Live Call Guidance, Competitive Battle Cards, and Playbooks and Coaching Rollout, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Nomi evals scored?+

Every case generated for Nomi — across Live Call Guidance and Competitive Battle Cards and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Nomi library include?+

The full Nomi library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Signal detection from live speech and Diagnosis quality under Live Call Guidance); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Nomi or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Nomi areas and set them up in a Corsac workspace, where you can run every test case against Nomi or your own agent with your own data.