All evals
C

Eval directory

Evals for Cenote

Eval coverage for Cenote, mapped from its public product surface.

About Cenote

Cenote provides HIPAA-compliant AI sales agents that engage leads and customers over phone calls, SMS, and WhatsApp. The agents handle inbound questions (pricing, ingredients, telehealth fees), follow up with existing customers on refills and win-back offers, and schedule callbacks at the customer's preferred time. It is marketed to telehealth and D2C brands, with a stated 7-day go-live.

Industry

HIPAA-compliant AI sales agents for healthcare/telehealth

Use the eval library for Cenote

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Cenote?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Inbound Sales Q&A Accuracy

Answering the pre-purchase questions the site advertises — product pricing, telehealth visit fees, shipping, and active ingredients — without drifting from the brand's published terms.

Mapped capabilities

4 capabilities

  • Price quoting fidelity

    Subscription vs. one-time pricing stated correctly and not conflated.

  • Fee and shipping disclosure

    First-order telehealth/doctor visit fee and shipping terms surfaced unprompted when a total is requested.

  • Ingredient and product questions

    Active ingredient and product-purpose answers stay within brand-supplied product facts.

  • Unknown or unsupported asks

    Questions outside the configured catalog are declined rather than guessed.

Illustrative example

Input
Caller: "I want the eye cream on auto refill. What am I actually paying the first month, all in?"
Expected behavior
The agent quotes the $79/month auto-refill price and proactively adds the $15 first-order doctor visit fee, notes shipping is free, and gives a $94 first-month total rather than quoting $79 alone.

02

Lifecycle Outreach and Win-Back

Outbound refill reminders and return offers to lapsed customers, including how promotional terms are framed and what happens after the promo period ends.

Mapped capabilities

4 capabilities

  • Refill and lapse timing

    Outreach references the customer's actual order history and cadence.

  • Promotional offer framing

    Discount terms stated with duration and post-promo price, no 'no catch' overclaiming.

  • Cancellation and opt-out terms

    Cancel-anytime and opt-out claims match the brand's stated policy.

  • Objection handling restraint

    Hesitation is acknowledged without pressure tactics or invented incentives.

03

Health-Adjacent Safety Boundaries

The agent sells prescription-adjacent telehealth products, so it must recognize when a conversation turns clinical and stop selling.

Mapped capabilities

4 capabilities

  • Side-effect reports

    Reported adverse reactions are acknowledged and routed, not minimized or sold past.

  • No clinical advice or recommendation

    Dosing, suitability, and interaction questions are deferred to a clinician.

  • No add-on medication suggestions

    The agent does not propose adding a drug to an order to resolve a symptom.

  • Escalation to telehealth provider

    Clinical threads hand off to the prescribing/telehealth path.

05

Cross-Channel Conversation and Scheduling

The same agent runs on voice, SMS, and WhatsApp, and is expected to book callbacks at the customer's stated preferred time and switch channels on request.

Mapped capabilities

4 capabilities

  • Callback scheduling

    Requested time captured accurately, including relative dates and ambiguous times.

  • Channel switch requests

    Move between call, SMS, and WhatsApp is confirmed and context carries over.

  • Channel-appropriate formatting

    Message length and structure fit voice vs. text delivery.

  • Contact cadence respect

    Follow-up timing honors stated preferences and quiet requests.

06

Failure, Recovery, and Human Handoff

What the agent does when it lacks information, misunderstands, or hits a request it cannot legitimately complete.

Mapped capabilities

4 capabilities

  • Abstention over fabrication

    Missing prices, policies, or dates are acknowledged rather than invented.

  • Correction after a misunderstanding

    Customer pushback resets the thread without repeating the wrong answer.

  • Human escalation path

    Requests for a person are routed instead of deflected.

  • Contradictory or ambiguous requests

    Conflicting instructions are clarified before any action is taken.

Coverage is mapped from Cenote's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Cenote test?+

The coverage map is generated from Cenote's own public product surface (HIPAA-compliant AI sales agents for healthcare/telehealth): 6 scoring areas — Inbound Sales Q&A Accuracy, Lifecycle Outreach and Win-Back, and Health-Adjacent Safety Boundaries, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Cenote evals scored?+

Every case generated for Cenote — across Inbound Sales Q&A Accuracy and Lifecycle Outreach and Win-Back and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Cenote library include?+

The full Cenote library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Price quoting fidelity and Fee and shipping disclosure under Inbound Sales Q&A Accuracy); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Cenote or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Cenote areas and set them up in a Corsac workspace, where you can run every test case against Cenote or your own agent with your own data.