All evals
PA

Eval directory

Evals for Peppr AI

Eval coverage for Peppr AI, mapped from its public product surface.

About Peppr AI

Peppr is an AI copilot that listens to live sales calls, detects the questions being asked, and surfaces short sourced answers drawn from a customer's connected knowledge sources. It also coaches reps against a configured playbook in real time, scores adherence, and writes notes and next steps after the call. It indexes across communication, knowledge base, CRM, support, dev, and storage tools, and can be run as a Peppr-hosted cloud service or as a single-tenant deployment in the customer's own environment.

Industry

real-time in-call sales intelligence / AI sales copilot

Use the eval library for Peppr AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Peppr AI?

6 scoring areas · 24 capabilities mapped · grounded in 7 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Live question detection

Recognizing that a question has been asked during a live call, across voice and chat, and surfacing a suggested response fast enough to be used before the rep has to stall.

Detects questions live across voice and chat www.usepeppr.ai

Mapped capabilities

4 capabilities

  • Question onset detection in speech

    Identifies the question the moment it lands in a live voice conversation, including questions embedded mid-sentence or asked as statements.

  • Chat-channel question detection

    Catches questions typed in the call's chat surface, not only spoken ones.

  • Non-question suppression

    Avoids firing on rhetorical asides, rep self-talk, and prospect statements that carry no ask.

  • Rep-invoked assist

    Handles explicit on-demand invocation (the Assist action) as a separate path from automatic detection.

02

Grounded answer surfacing

Producing short answers drawn from the customer's connected knowledge sources, with the source shown, and declining when the connected corpus does not support an answer.

Peppr listens to your calls and surfaces the right answer before your customer finishes asking. www.usepeppr.ai

Mapped capabilities

4 capabilities

  • Retrieval from connected sources

    Answers are pulled from the knowledge sources the customer has connected, not from general model knowledge.

  • Source attribution

    Each surfaced answer carries the source it was drawn from so the rep can check it.

  • Abstention on uncovered questions

    When the corpus has no supporting content, the system declines rather than producing an unsourced answer.

  • Answer brevity for live use

    Responses stay short enough to be read and spoken while the call is still running.

Illustrative example

Input
Live call, prospect asks: "What's your uptime SLA for the single-tenant deployment?" The customer's connected knowledge sources contain no SLA or uptime commitment for single-tenant.
Expected behavior
Peppr surfaces no answer, or surfaces an explicit no-coverage state, and does not state an uptime figure. The unanswered question is routed as a knowledge gap after the call.

03

Playbook coaching and adherence

Keeping reps on a configured methodology during the call through live nudges, and scoring how well each call followed it.

Adherence scored on every call www.usepeppr.ai

Mapped capabilities

4 capabilities

  • Configured playbook execution

    Coaches against the methodology the customer configured (for example MEDDIC or SPICED) rather than a fixed default.

  • Stage and item progression

    Tracks which playbook items have been covered and which stage the call is in as the conversation moves.

  • Live nudges

    Prompts the rep toward uncovered playbook items while the call is in progress.

  • Adherence scoring

    Produces an adherence score per call that a manager can use to see where to step in.

Illustrative example

Input
Playbook configured as MEDDIC. Rep completes Metrics and Identify Pain in the Qualify stage but never asks who signs. Call ends.
Expected behavior
Qualify shows two of three items covered and the economic buyer item under Decision stays uncovered. The adherence score reflects the missing item, and the post-call output flags economic buyer as not established.

04

Post-call output

What Peppr writes after the call ends: notes, next steps, question classification, and routing of knowledge the corpus could not answer.

Notes and next steps, written automatically www.usepeppr.ai

Mapped capabilities

4 capabilities

  • Automatic call notes

    Generates notes from the call without the rep authoring them.

  • Next steps extraction

    Pulls the commitments and follow-ups agreed on the call into explicit next steps.

  • Question classification

    Classifies the questions asked during the call for downstream analysis.

  • Knowledge gap routing

    Identifies questions the connected sources could not answer and routes them as gaps.

05

Connector coverage and indexing

Indexing across the customer's stack — communication, knowledge base, project management, code and dev, CRM and sales, support, and storage — so answers surface wherever content lives.

Pulls answers straight from your knowledge base www.usepeppr.ai

Mapped capabilities

4 capabilities

  • Cross-category retrieval

    Draws on all connected categories rather than favoring a single system of record.

  • Per-connector authentication and scope

    Connects each source under its own credentials and respects the scope granted.

  • Index freshness

    Reflects updates to connected sources so answers do not go stale against current documentation.

  • Unconnected-source behavior

    Behaves predictably when a relevant tool in the stack has not been connected.

06

Deployment and data boundaries

Honoring the difference between the Peppr-hosted cloud offering and a single-tenant deployment in the customer's own environment, including the controller/processor split described in the privacy policy.

Peppr acts as a processor that handles that data on the Customer's behalf and under their instructions www.usepeppr.ai

Mapped capabilities

4 capabilities

  • Cloud vs. single-tenant behavior

    Data handling follows the deployment model set out in the order form rather than one averaged behavior.

  • Customer-instructed processing

    Acts as a processor on customer content under the customer's instructions in customer-controlled deployments.

  • Tenant content isolation

    Content connected by one customer does not surface in another customer's answers.

  • Screen-side invisibility

    Assist runs on the rep's screen without exposing itself to the prospect on the call.

Coverage is mapped from Peppr AI's public pages (7 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Peppr AI test?+

The coverage map is generated from Peppr AI's own public product surface (real-time in-call sales intelligence / AI sales copilot): 6 scoring areas — Live question detection, Grounded answer surfacing, and Playbook coaching and adherence, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Peppr AI evals scored?+

Every case generated for Peppr AI — across Live question detection and Grounded answer surfacing and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Peppr AI library include?+

The full Peppr AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Question onset detection in speech and Chat-channel question detection under Live question detection); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Peppr AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Peppr AI areas and set them up in a Corsac workspace, where you can run every test case against Peppr AI or your own agent with your own data.