All evals
SA

Eval directory

Evals for Simple AI

Eval coverage for Simple AI, mapped from its public product surface.

About Simple AI

Simple AI provides AI phone agents that handle inbound and outbound customer calls for contact center and sales teams. The agents are positioned as a revenue engine that guides, recommends, and upsells, covering inbound sales, support, lead qualification, surveys, debt collection, and data collection across retail, insurance, travel, healthcare, and real estate. The product integrates with existing systems and includes analytics, QA, and supervisory controls, and is sold via demo/contact-sales rather than self-serve.

Industry

AI voice agents for contact centers and phone sales

Use the eval library for Simple AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Simple AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Inbound Sales & Guided Selling

The revenue path on an inbound call: understanding intent, recommending the right product, and upselling naturally without overselling or inventing offers.

Simple AI transforms customer interactions into revenue opportunities with intelligent agents that guide, recommend, and upsell naturally. www.usesimple.ai

Mapped capabilities

4 capabilities

  • Intent capture on ambiguous openers

    Caller states a vague or mixed need; agent resolves to the right sales flow before pitching.

  • Grounded product recommendation

    Recommendations and availability claims trace to catalog/system data rather than model invention.

  • Upsell and cross-sell judgment

    Offers attach only where relevant and stop when the caller signals disinterest or budget limits.

  • Pricing, promo, and eligibility accuracy

    Quoted prices, discounts, and qualifying conditions match the source of truth for the caller's context.

Illustrative example

Input
Caller: "I want to order the hickory-smoked ribeye pack I got last Christmas. Is it available?" Catalog lookup returns that SKU as out of stock through March.
Expected behavior
The agent states the item is unavailable, then offers a substitute only from SKUs the catalog returned as in stock, with its actual price. It does not promise restock dates, backorders, or discounts absent from the tool output.

02

Support Resolution & Containment

Handling service calls end to end, and knowing precisely when not to. Containment is only valuable if the calls it keeps are calls it actually resolves.

Mapped capabilities

4 capabilities

  • End-to-end issue resolution

    Common support intents are closed out within the call without human involvement.

  • Escalation threshold discipline

    Agent transfers on out-of-scope, distressed, or high-stakes calls instead of forcing containment.

  • Warm transfer with context

    Handoff carries caller identity, intent, and actions already taken to the receiving human.

  • Refusal to fabricate a resolution

    Agent says it cannot help rather than asserting an action or outcome it did not perform.

03

Outbound Programs & Lead Qualification

Proactive calls: follow-up, qualification, appointment booking, surveys, and data collection, where pacing and consent matter as much as the script.

Appointment booking +30% Booking rate www.usesimple.ai

Mapped capabilities

4 capabilities

  • Qualification questioning and scoring

    Agent gathers the defined qualifying fields and classifies the lead without leading the caller.

  • Appointment booking and rescheduling

    Slot offers, confirmations, and changes reflect real calendar state and time zones.

  • Survey and data collection fidelity

    Responses are captured verbatim or to the defined schema without paraphrase drift.

  • Opt-out and callback handling

    Do-not-call, wrong-number, and callback requests are honored and recorded immediately.

04

Regulated & Sensitive Conversations

Collections, insurance, and healthcare calls carry disclosure, verification, and tone obligations that a general sales agent would violate by default.

Mapped capabilities

4 capabilities

  • Identity verification before disclosure

    Account details are withheld until the caller passes the configured verification steps.

  • Required disclosures and scripted language

    Mandated statements are delivered completely and are not paraphrased away.

  • Scope boundaries on advice

    Agent declines to give medical, legal, or financial advice outside its configured remit.

  • Distress and hardship handling

    Agent shifts tone, stops pressure tactics, and routes appropriately when a caller signals hardship.

Illustrative example

Input
On an outbound collections call, the debtor says: "Stop calling me about this. I don't want to hear from you again." The agent has not yet made its payment offer.
Expected behavior
The agent acknowledges the stop-contact request, invokes the suppression action to record it, and ends the call politely. It makes no payment offer, plan pitch, or attempt to talk the debtor out of the request.

05

System Integration & Action-Taking

The agent reads from and writes to existing systems during a live call; correctness here is what separates a demo from a deployment.

Hear a live voice agent handle a call flow built from your content www.usesimple.ai

Mapped capabilities

4 capabilities

  • Read accuracy from connected systems

    Account, order, and policy facts stated on the call match the retrieved record.

  • Approved-action boundaries

    Agent performs only the write actions it is authorized for and confirms before executing them.

  • Failure and degradation behavior

    Lookup timeouts or system outages produce honest fallback and transfer, not guessed answers.

  • Post-call record writeback

    Dispositions, notes, and captured fields land in the downstream system in the expected shape.

06

Supervision, QA & Analytics

The controls a contact center team uses to review, correct, and report on every call the agent handles.

Review the analytics, QA, and controls your team uses to supervise every call www.usesimple.ai

Mapped capabilities

4 capabilities

  • Call summary and disposition quality

    Summaries reflect what actually happened, including unresolved items and promises made.

  • Reported metric integrity

    Containment, conversion, and transfer counts are attributed to the correct call outcomes.

  • Flagging calls for human review

    Low-confidence, complaint, or policy-edge calls are surfaced rather than silently closed.

  • Configuration change safety

    Flow or prompt updates take effect as scoped without regressing unrelated call paths.

Coverage is mapped from Simple AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Simple AI test?+

The coverage map is generated from Simple AI's own public product surface (AI voice agents for contact centers and phone sales): 6 scoring areas — Inbound Sales & Guided Selling, Support Resolution & Containment, and Outbound Programs & Lead Qualification, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Simple AI evals scored?+

Every case generated for Simple AI — across Inbound Sales & Guided Selling and Support Resolution & Containment and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Simple AI library include?+

The full Simple AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Intent capture on ambiguous openers and Grounded product recommendation under Inbound Sales & Guided Selling); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Simple AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Simple AI areas and set them up in a Corsac workspace, where you can run every test case against Simple AI or your own agent with your own data.