All evals
SA

Eval directory

Evals for Siena AI

Eval coverage for Siena AI, mapped from its public product surface.

About Siena AI

Siena is an AI customer experience platform for consumer and ecommerce brands, positioned as an "AI CX operating system" with empathic AI agents that work across customer channels. Its products include a CX Agent, a Reviews Agent, Ask Siena for brand operators, and Siena Docs, a knowledge layer that supplies agents with brand policies and context. Pricing is quote-based, structured around a monthly platform fee plus a per-automated-ticket rate.

Industry

ecommerce customer service AI agents

Use the eval library for Siena AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Siena AI?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

CX Agent conversation handling

Core customer-facing resolution quality across the channels and message types the CX Agent covers, including multi-turn ecommerce support conversations and video-based inputs for warranty, troubleshooting, and product questions.

One intelligence layer powering agents across every customer surface. www.siena.cx

Mapped capabilities

4 capabilities

  • Order-context resolution

    Answering questions about a specific order using retrieved order state rather than generic guidance.

  • Multi-turn intent tracking

    Holding the customer's goal and prior commitments across turns without re-asking for details already given.

  • Video and multimodal intake

    Handling customer-submitted video for warranty claims, troubleshooting, and product questions.

  • Escalation and handoff to humans

    Recognizing conversations that need a human and routing them instead of improvising.

02

Policy fidelity and eligibility decisions

Whether the agent applies brand refund, return, exchange, and warranty policy exactly as written — including windows, condition requirements, and exclusions — rather than approximating or inventing terms.

Send warranty claims, troubleshooting issues, or product questions as video. www.siena.cx

Mapped capabilities

4 capabilities

  • Return window arithmetic

    Correctly comparing order or ship dates against the stated eligibility window.

  • Exclusion handling

    Respecting carve-outs such as final-sale, personalized, worn, or washed items.

  • Refund timing claims

    Stating processing timelines only as documented, without optimistic rounding.

  • Price matching rules

    Applying post-purchase price adjustment logic within its documented bounds.

Illustrative example

Input
Customer writes: "I bought these clearance leggings 12 days ago and want to return them. Your policy says 30 days, right?" Docs list a 30-day window with final-sale items excluded.
Expected behavior
The agent recognizes the order is inside the 30-day window but that the clearance item is final sale, declines the refund on that ground, and cites the exclusion. It does not approve the return by applying the window alone.

03

Agentic actions and guardrails

The write-side surface: issuing refunds, applying discount codes, generating return labels, and adding items to a cart. Covers whether the agent takes the authorized action, stops at the boundary of its authority, and confirms before irreversible steps.

Mapped capabilities

4 capabilities

  • Authorized action execution

    Completing refunds, label generation, and code application when eligibility is met.

  • Action boundary refusal

    Declining actions outside granted authority instead of claiming they were done.

  • Confirmation before irreversible steps

    Seeking customer assent before committing a refund or exchange path.

  • Proactive offers and upsell restraint

    Making recommendations only where warranted and never at the cost of resolving the request.

04

Siena Docs knowledge grounding

The knowledge layer that supplies agents with brand policies and context, including verified versus in-review status, document freshness, and conflicting versions of the same policy.

Siena Docs is one home for everything your team knows. www.siena.cx

Mapped capabilities

4 capabilities

  • Citation to source document

    Grounding policy answers in the governing doc rather than model priors.

  • Stale and conflicting policy detection

    Preferring the live version when multiple versions of a policy exist.

  • Verified vs. in-review content

    Treating unverified or pending pages with appropriate caution.

  • Structured product data lookup

    Reading sizing, color, SKU, and restock tables accurately.

Illustrative example

Input
Knowledge base contains two return-policy pages: a verified one updated 2 days ago stating 30 days, and an older unverified page stating 60 days. Customer asks how long they have to return.
Expected behavior
The agent answers 30 days, sourcing the verified current page, and does not surface or hedge with the 60-day figure from the outdated unverified document.

05

Brand voice and empathic consistency

Siena positions itself on stable, consistently on-brand output driven by a tone-of-voice playbook. This area covers adherence to documented voice rules and appropriate handling of frustrated or emotionally charged customers.

Mapped capabilities

3 capabilities

  • Tone playbook adherence

    Following documented rules such as leading with the solution.

  • Consistency across channels

    Holding the same voice whether the surface is chat, email, or reviews.

  • Empathy under complaint pressure

    Staying warm and non-defensive when a customer is upset.

06

Reviews Agent and Ask Siena insight

The non-inbox surfaces: end-to-end review management for public customer reviews, and Ask Siena's operator-facing synthesis of what customers are actually saying across conversations.

Meet the first AI agent built specifically for end-to-end review management. www.siena.cx

Mapped capabilities

4 capabilities

  • Review response appropriateness

    Matching response tone and substance to review sentiment and rating.

  • Operator query synthesis

    Answering brand-operator questions with themes drawn from real conversation data.

  • Insight faithfulness to evidence

    Reporting only patterns the underlying conversations support, without embellishment.

  • Recurring report generation

    Producing periodic voice-of-customer summaries for internal stakeholders.

Coverage is mapped from Siena AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Siena AI test?+

The coverage map is generated from Siena AI's own public product surface (ecommerce customer service AI agents): 6 scoring areas — CX Agent conversation handling, Policy fidelity and eligibility decisions, and Agentic actions and guardrails, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Siena AI evals scored?+

Every case generated for Siena AI — across CX Agent conversation handling and Policy fidelity and eligibility decisions and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Siena AI library include?+

The full Siena AI library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Order-context resolution and Multi-turn intent tracking under CX Agent conversation handling); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Siena AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Siena AI areas and set them up in a Corsac workspace, where you can run every test case against Siena AI or your own agent with your own data.