All evals
Decagon

Eval directory · Customer Support

Evals for Decagon

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Decagon AI products.

About Decagon

Decagon builds AI customer support agents that understand full conversation context, integrate with existing helpdesks, and resolve tickets end-to-end without human intervention. Its platform is used by fintechs, SaaS companies, and consumer platforms.

Employees

~80

Industry

Customer Support AI

Headquarters

San Francisco, CA

Website

decagon.ai

Use the eval library for Decagon

All 241 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Related in Customer Support

All evals →

More Customer Support eval libraries

Coverage map

What would you measure for Decagon?

4 areas · 241 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Conversational Quality Resolution Accuracy

Mapped capabilities

61 scenarios

  • Single-turn FAQ exact-match resolution

Public sample case

Input
The enterprise operator's KB article 'Return Policy v3' states: 'Customers may return eligible items within 30 days of delivery. This policy applies exclusively to purchases made on or after 2024-01-01. Purchases made before 2024…
Expected behavior
The agent's response explicitly states both (1) the 30-day return window and (2) the condition that the policy applies only to purchases made on or after 2024-01-01. Neither clause may be omitted or softened. If the agent cannot determine the customer's purchase date from context, it must ask befor…
Check
Pass / fail check

02

Deflection Vs Escalation Decisioning

Mapped capabilities

66 scenarios

  • Sentiment Rising-Anger Escalation Trigger

Public sample case

Input
A customer contacts support about a duplicate charge. Over four consecutive turns their language escalates: Turn 1 is mildly annoyed ('this seems wrong'), Turn 2 expresses clear frustration ('I've already checked this twice'), Tu…
Expected behavior
The agent detects the upward-trending sentiment delta across the four-turn window — not just turn-4 in isolation — and initiates escalation to a human agent. The warm handoff package cites the trajectory (e.g., 'sentiment rose across turns 1–4; no single message individually crossed threshold') rat…
Check
Pass / fail check

03

Tool Use Precision Recall Against Connected Systems

Mapped capabilities

60 scenarios

  • Correct tool selection for single-system intents

Public sample case

Input
Tenant is a D2C apparel brand running Shopify storefront with Stripe as payment processor. Both shopify_create_refund and stripe_create_refund tools are registered. Order #84231 was placed through Shopify checkout but payment was…
Expected behavior
Agent retrieves the order to obtain the Stripe payment_intent_id, then calls stripe_create_refund with that payment_intent_id and the correct amount. It does not call shopify_create_refund at any point. It confirms the refund amount and expected posting timeline to the customer.
Check
Pass / fail check

04

Transactional High Risk Action Safety

Mapped capabilities

54 scenarios

  • Refund within-policy auto-approval

Frequently asked questions

What do the Corsac evals for Decagon test?+

Each eval pack tests Decagon's public product surface — including Conversational Quality Resolution Accuracy, Deflection Vs Escalation Decisioning, and Tool Use Precision Recall Against Connected Systems — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Decagon evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 241 Decagon cases — from Deflection Vs Escalation Decisioning (66 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Decagon library.

How many test cases does the Decagon library include?+

The Decagon eval library includes 241 graded test cases across 4 eval packs, the largest being Deflection Vs Escalation Decisioning with 66 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Decagon or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 4 Decagon packs — Conversational Quality Resolution Accuracy and Deflection Vs Escalation Decisioning and the rest — against Decagon or your own agent with your own data.