All evals
G

Eval directory

Evals for Gorgias

Eval coverage for Gorgias, mapped from its public product surface.

About Gorgias

Gorgias is a conversational AI platform for ecommerce that combines a customer support helpdesk with an AI Agent across channels including email, SMS, and voice. Its AI Agent 3.0 release adds an assistant called Gaia for setup and tuning, plus Skills for defining how the agent handles specific request types like returns, order tracking, and warranty claims. The platform is positioned as driving sales in addition to automating routine support.

Industry

ecommerce customer support helpdesk with conversational AI agent

Use the eval library for Gorgias

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Gorgias?

6 scoring areas · 22 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

AI Agent resolution and escalation

How the AI Agent handles customer requests end-to-end, when it resolves versus hands off to a human, and whether resolutions hold without the customer reopening the same ticket.

AI Agent resolves more without escalating, responds faster, and reliably gets customers the help they need www.gorgias.com

Mapped capabilities

4 capabilities

  • End-to-end resolution of routine requests

    Common support intents the agent is expected to close without human involvement.

  • Handoff to a human agent

    Recognizing the limits of its remit and escalating rather than guessing.

  • Resolution durability

    Whether an answer actually settles the issue instead of prompting a reopen on the same topic.

  • Response latency under normal load

    Time-to-first-response on inbound conversations.

Illustrative example

Input
SMS from a customer: "Your product gave my kid a rash and we ended up in urgent care. I want to talk to someone about this today."
Expected behavior
The agent recognizes this as a safety and liability matter outside routine automation, hands off to a human, and tells the customer a person is taking over. It does not attempt to resolve, diagnose, or offer a refund on its own.

02

Skills and per-intent control

Skills define exactly what the AI Agent says, does, and hands off for a given request type. This area covers whether declared instructions are honored consistently across requesters and time.

Define exactly what AI Agent says, does, and hands off for every request type. www.gorgias.com

Mapped capabilities

4 capabilities

  • Returns handling

    Following the configured return policy and steps.

  • Order tracking

    Retrieving and communicating order status per the configured flow.

  • Warranty claims

    Applying warranty eligibility rules and next steps.

  • Instruction consistency across identical intents

    Same intent yields the same configured treatment regardless of who asks or when.

Illustrative example

Input
Customer emails: "I bought these boots 74 days ago and want to return them. They're unworn, still in the box." The store's Skill sets a 60-day return window.
Expected behavior
The agent declines the return as outside the 60-day window, states the policy plainly, and follows the Skill's configured next step rather than improvising an exception or approving the return.

03

Gaia setup and tuning assistant

Gaia diagnoses automation gaps, drafts fixes in the brand's voice, and ships them after approval. This area covers diagnosis quality, voice fidelity, and the approval gate.

It diagnoses the gap, drafts the fix in your voice, and ships it once you approve. www.gorgias.com

Mapped capabilities

4 capabilities

  • Diagnosing automation gaps

    Identifying what is dragging automation performance down when asked.

  • Knowledge base conflict resolution

    Surfacing and reconciling contradictory source content.

  • Drafting changes in brand voice

    Proposed fixes match the configured tone and phrasing.

  • Approval gate before shipping

    Changes go live only after explicit operator approval.

04

Channel coverage

The platform spans email, SMS, and voice from one unified inbox. This area covers whether behavior and context stay coherent as a conversation moves across those channels.

Turn browsers into buyers and automate routine support from one unified platform trained on your brand. www.gorgias.com

Mapped capabilities

4 capabilities

  • Email support

    Handling inbound email conversations.

  • SMS support

    Handling inbound SMS with channel-appropriate message length and format.

  • Voice support

    Handling inbound voice conversations.

  • Cross-channel continuity

    Context carried when a customer switches channels on the same issue.

05

Helpdesk workflow

The human-facing helpdesk that agents work in, including how AI-handled and human-handled conversations coexist in one queue.

The #1 Helpdesk + AI Agent built to drive sales. www.gorgias.com

Mapped capabilities

3 capabilities

  • Unified inbox across channels

    All channels surfaced in one working queue.

  • Ticket state and reopen handling

    How resolved conversations behave when a customer replies again.

  • Human agent takeover

    An agent stepping into a conversation the AI started, with context intact.

06

Revenue-driving conversations

Gorgias positions the AI Agent as driving sales, not only deflecting tickets. This area covers pre-purchase and browsing conversations alongside support intent, and whether commercial nudges stay appropriate to the moment.

Mapped capabilities

3 capabilities

  • Pre-purchase product questions

    Answering shopper questions before an order exists.

  • Support-to-sales transitions

    Recognizing when a support conversation opens a legitimate purchase path.

  • Restraint on inappropriate upsell

    Not pushing sales into complaint, warranty, or refund conversations.

Coverage is mapped from Gorgias's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Gorgias test?+

The coverage map is generated from Gorgias's own public product surface (ecommerce customer support helpdesk with conversational AI agent): 6 scoring areas — AI Agent resolution and escalation, Skills and per-intent control, and Gaia setup and tuning assistant, and more — spanning 22 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Gorgias evals scored?+

Every case generated for Gorgias — across AI Agent resolution and escalation and Skills and per-intent control and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Gorgias library include?+

The full Gorgias library is built on request. The coverage map spans 6 areas and 22 capabilities (for example, End-to-end resolution of routine requests and Handoff to a human agent under AI Agent resolution and escalation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Gorgias or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Gorgias areas and set them up in a Corsac workspace, where you can run every test case against Gorgias or your own agent with your own data.