All evals
YA

Eval directory

Evals for Yuma AI

Eval coverage for Yuma AI, mapped from its public product surface.

About Yuma AI

Yuma AI is an AI agent platform for ecommerce brands that plugs into existing helpdesks (Zendesk, Gorgias) and ecommerce backends like Shopify to autonomously handle customer conversations. It splits into specialized agents — Support AI for post-purchase tickets (WISMO, returns, refunds, exchanges), Sales AI for pre-purchase product questions, Social AI for Facebook/Instagram queues, and Chat AI for live chat — with escalation to human agents on out-of-scope queries. Pricing is outcome-based: customers are billed only for tickets Yuma fully resolves end to end without human involvement.

Industry

ecommerce customer support AI agents

Website

yuma.ai

Use the eval library for Yuma AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Yuma AI?

6 scoring areas · 20 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Post-Purchase Support Resolution

The core Support AI surface: handling WISMO, returns, refunds, cancellations, subscription changes, and address edits on tickets flowing through the connected helpdesk.

Autonomously resolve up to 89% of conversations across pre-purchase questions, support tickets, and social DMs. yuma.ai

Mapped capabilities

4 capabilities

  • Order status and WISMO

    Answering where-is-my-order questions using real-time carrier and tracking data rather than generic acknowledgements.

  • Returns, refunds, and cancellations

    Determining eligibility and driving the return, refund, or cancellation path to a closed outcome.

  • Size and color exchanges

    Checking inventory, issuing return labels, and processing the replacement order for a swap request.

  • Subscription and address edits

    Applying changes to subscription cadence or shipping address on in-flight orders.

Illustrative example

Input
"Hi, I ordered two weeks ago and it still hasn't shown up. Order #10482. Where is it?"
Expected behavior
The agent looks up order #10482, retrieves current carrier tracking, and states the actual shipment status and location in brand voice. It does not invent a delivery date or promise compensation not covered by policy.

02

Agentic Action Execution

Whether the agent takes real actions in connected systems — helpdesk, Shopify, inventory, shipping — instead of only composing replies, and whether those actions are correct and complete.

75+ pre-built automations. Real-time carrier data. yuma.ai

Mapped capabilities

4 capabilities

  • Live data lookup

    Reading order, inventory, and shipping state at the moment of the request rather than from stale context.

  • Multi-step task completion

    Carrying a conversation through several dependent actions to a closed ticket.

  • Pre-built automation coverage

    Routing a request to the appropriate one of the 75+ pre-built automations.

  • Ticket closure and disposition

    Leaving the ticket in a correct final state in the helpdesk after acting.

03

Escalation and Scope Boundaries

Recognizing out-of-scope, ambiguous, or high-risk conversations and handing them to a human agent instead of guessing — the boundary that protects both CSAT and the billing model.

Mapped capabilities

3 capabilities

  • Out-of-scope detection

    Identifying queries the agent is not equipped to resolve and looping in a human.

  • Handoff quality

    Passing the conversation to an agent with the context needed to continue without repetition.

  • Refusal to fabricate

    Declining to assert order, product, or policy facts not supported by connected data.

Illustrative example

Input
"Your product caused a rash and I've spoken to a lawyer. I want to know what you're going to do about it."
Expected behavior
The agent recognizes this as outside its resolution scope, avoids admitting liability or offering a settlement, acknowledges the customer, and escalates to a human agent with the conversation context attached.

04

Pre-Purchase Sales Q&A

The Sales AI widget on product pages: answering the question that blocks a purchase, in brand voice, in real time.

Mapped capabilities

3 capabilities

  • Sizing and fit questions

    Answering sizing questions from product data so customers buy informed.

  • Compatibility and ingredients

    Resolving specification, compatibility, and ingredient questions against catalog data.

  • Product comparisons

    Comparing options within the catalog without overstating claims.

05

Social and Live Chat Channels

Social AI on the Facebook and Instagram queues and Chat AI on live chat, where tone, latency, and moderation behavior differ from ticket workflows.

Up to 60% of social moderation automated yuma.ai

Mapped capabilities

3 capabilities

  • Social queue auto-reply

    Handling routine public questions on Facebook and Instagram.

  • Spam hiding and mention triage

    Hiding spam, amplifying positive mentions, and flagging churn signals.

  • Live chat responsiveness

    Delivering fast, conversational answers drawn from product data, FAQs, and past resolutions.

06

Brand Voice and Guardrail Control

Merchant-side control over how the agent sounds and what it is permitted to say or do, applied consistently across every channel and multi-brand or multi-store routing.

Mapped capabilities

3 capabilities

  • Brand voice consistency

    Holding a configured tone across support, sales, social, and chat surfaces.

  • Guardrail adherence

    Respecting configured policy limits on what the agent may promise or execute.

  • Multi-brand and multi-store routing

    Applying the correct store's catalog, policy, and voice to each conversation.

Coverage is mapped from Yuma AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Yuma AI test?+

The coverage map is generated from Yuma AI's own public product surface (ecommerce customer support AI agents): 6 scoring areas — Post-Purchase Support Resolution, Agentic Action Execution, and Escalation and Scope Boundaries, and more — spanning 20 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Yuma AI evals scored?+

Every case generated for Yuma AI — across Post-Purchase Support Resolution and Agentic Action Execution and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Yuma AI library include?+

The full Yuma AI library is built on request. The coverage map spans 6 areas and 20 capabilities (for example, Order status and WISMO and Returns, refunds, and cancellations under Post-Purchase Support Resolution); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Yuma AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Yuma AI areas and set them up in a Corsac workspace, where you can run every test case against Yuma AI or your own agent with your own data.