All evals
Sierra AI

Eval directory · Customer Support

Evals for Sierra AI

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Sierra AI AI products.

About Sierra AI

Sierra AI builds conversational AI agents for customer experience, designed to handle the full resolution lifecycle across every channel — chat, voice, and messaging. Sierra agents are deployed by leading consumer brands to reduce handle time and improve CSAT.

Employees

~200

Industry

Customer Experience AI

Headquarters

San Francisco, CA

Website

sierra.ai

Use the eval library for Sierra AI

All 70 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Related in Customer Support

All evals →

More Customer Support eval libraries

Coverage map

What would you measure for Sierra AI?

7 areas · 70 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Brand Voice Policy Guardrails

Evaluates Sierra's Brand Voice & Policy Guardrails across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Enterprise conversational AI agents eval coverage.

Mapped capabilities

10 scenarios

  • Ghostwriter brand warmth bound
  • PCI refusal full card in chat
  • Competitor disparagement guardrail

02

Connected Systems Tool Audit

Evaluates Sierra's Connected Systems & Tool Audit across 11 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Enterprise conversational AI agents eval coverage.

Mapped capabilities

11 scenarios

  • Zendesk ticket append with audit
  • Shopify partial refund guard
  • Tool trace completeness for Explorer

03

Escalation Live Assist Handoff

Evaluates Sierra's Escalation & Live Assist Handoff across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Enterprise conversational AI agents eval coverage.

Mapped capabilities

10 scenarios

  • Warm handoff package completeness
  • Self-service deflection inappropriate for legal threat
  • Live Assist whisper mode privacy

04

Experiments Observability Safety

Evaluates Sierra's Experiments & Observability Safety across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Enterprise conversational AI agents eval coverage.

Mapped capabilities

9 scenarios

  • Experiment variant assignment stickiness
  • Unsafe variant auto-rollback
  • Explorer PII scrub on export

05

Knowledge Grounding Agent Memory

Evaluates Sierra's Knowledge Grounding & Agent Memory across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Enterprise conversational AI agents eval coverage.

Mapped capabilities

10 scenarios

  • KB citation required for policy answer
  • Agent memory personalization limit
  • Hallucinated API feature refusal

06

Skills Intent Routing

Evaluates Sierra's Skills & Intent Routing across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Enterprise conversational AI agents eval coverage.

Mapped capabilities

10 scenarios

  • Composable skill activation
  • Ambiguous dual-intent disambiguation
  • Skill precedence under guardrail block

07

Voice Channel Asr Safety

Evaluates Sierra's Voice Channel & ASR Safety across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Enterprise conversational AI agents eval coverage.

Mapped capabilities

10 scenarios

  • DTMF IVR payment path
  • Barge-in handling mid-prompt
  • ASR homophone repair account number

Frequently asked questions

What do the Corsac evals for Sierra AI test?+

Each eval pack tests Sierra AI's public product surface — including Brand Voice Policy Guardrails, Connected Systems Tool Audit, Escalation Live Assist Handoff — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Sierra AI evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Sierra AI library include?+

The Sierra AI eval library includes 70 graded test cases across 7 eval packs. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Sierra AI or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against Sierra AI or your own agent with your own data.