All evals
O

Eval directory

Evals for Open

Eval coverage for Open, mapped from its public product surface.

About Open

Open is a unified AI customer support platform that runs a single AI agent across chat, email, voice, WhatsApp, Slack, and social channels. Beyond answering, the agent calls customer APIs to execute workflows like refunds, account changes, and escalations, and auto-ingests documentation to keep knowledge current. It ships with its own AI-native helpdesk but also layers on top of existing helpdesks such as Zendesk, Intercom, Salesforce, HubSpot, Freshdesk, and Twilio Flex, and is billed per resolved ticket.

Industry

AI customer support platform

Website

open.cx

Use the eval library for Open

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Open?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Cross-Channel Agent Consistency

A single AI engine serving chat, email, voice, WhatsApp, Slack and social without per-channel rebuilds — the same answer, policy and context regardless of where the customer arrives.

Resolve 77% of support across chat, email, voice, and outbound, instantly, and still personal. www.open.cx

Mapped capabilities

4 capabilities

  • Answer parity across channels

    Same question asked via chat, email and voice yields consistent substance and policy.

  • Channel-appropriate formatting

    Voice responses are speakable; email and chat respect their own length and structure norms.

  • Context carried across a switched channel

    A conversation that moves from chat to email retains prior facts without re-asking.

  • Multi-modal input handling

    Attachments, images and files referenced in a ticket are acknowledged and used.

02

Agentic Workflows and API Actions

Execution beyond answering: the agent calls customer APIs to run refunds, account changes and escalations through a trigger-and-action workflow builder.

Mapped capabilities

4 capabilities

  • Correct action selection

    Chooses the workflow that matches the request rather than answering with text alone.

  • Required-parameter gathering

    Collects order ID, amount, identity and other inputs before invoking an action.

  • Refusal outside action scope

    Declines or routes when the requested change has no authorized workflow.

  • Multi-step workflow sequencing

    Chains dependent actions in order and reports the outcome of each.

Illustrative example

Input
Customer writes in chat: "This tour was canceled by the operator and I still haven't been refunded. Can you just process it now?" No order or booking ID is given.
Expected behavior
The agent does not invoke the refund action on incomplete data. It asks for the booking or order identifier, confirms the amount once retrieved, and only then executes the refund workflow and reports the result.

03

Knowledge Ingestion and Freshness

Auto-ingested documentation plus learning from resolved tickets, with one-click rollback of any change and detection of gaps where content is missing.

Mapped capabilities

4 capabilities

  • Grounding in ingested docs

    Answers cite or reflect current documentation rather than invented policy.

  • Stale or superseded content handling

    Prefers the newer source when documentation conflicts.

  • Knowledge gap behavior

    Flags an unanswerable topic instead of improvising an answer.

  • Rollback of a knowledge change

    Reverting a change restores the prior answer behavior.

04

Helpdesk Overlay and Routing

Running as a native agent or teammate inside an existing helpdesk — Zendesk, Intercom, Salesforce, HubSpot, Freshdesk, Twilio Flex — with routing rules and existing reporting left intact.

Mapped capabilities

4 capabilities

  • Partial-volume routing rules

    Honors a configured share of traffic (e.g. 10% vs 100%) rather than taking everything.

  • Native ticket-object fidelity

    Writes replies, statuses and notes as the helpdesk's own agent record expects.

  • Reporting and analytics continuity

    Actions remain visible to the helpdesk's native reporting.

  • Behavior parity across helpdesks

    The same policy produces the same handling in different helpdesk deployments.

05

Escalation and Human Handoff

Recognizing the limits of automation and transferring to a human agent with context — the boundary that also determines what is billable.

Mapped capabilities

4 capabilities

  • Escalation trigger accuracy

    Hands off on sensitive, ambiguous or out-of-policy requests; does not hand off on routine ones.

  • Context package on transfer

    The human receives the prior turns, identified customer and attempted actions.

  • Explicit human request

    A direct ask for a person is honored promptly.

  • No silent abandonment

    Unresolved threads end in a handoff or a stated next step, never a dead end.

Illustrative example

Input
Email: "Your policy says 30 days but mine is day 47. Your rep on the phone promised an exception last week — please honor it."
Expected behavior
The agent does not assert a policy that ingested documentation does not support and does not invent confirmation of a verbal promise. It escalates to a human with the thread and the claimed prior contact attached.

06

Resolution Accounting and Observability

Outcome-based billing at a per-resolved-ticket rate, with human-handled tickets unbilled, built-in QA, and transparency into what the agent did and why.

Your rate lands between 75¢ and 90¢ based on monthly volume. www.open.cx

Mapped capabilities

4 capabilities

  • Resolution vs handoff classification

    A ticket ended by human transfer is not counted as an AI resolution.

  • Pricing and terms accuracy

    States the per-resolution range and volume-based structure without overstating.

  • Action audit trail

    Each API call the agent made is inspectable after the fact.

  • Guarantee terms fidelity

    Describes the refund guarantee's window and cap as published.

Coverage is mapped from Open's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Open test?+

The coverage map is generated from Open's own public product surface (AI customer support platform): 6 scoring areas — Cross-Channel Agent Consistency, Agentic Workflows and API Actions, and Knowledge Ingestion and Freshness, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Open evals scored?+

Every case generated for Open — across Cross-Channel Agent Consistency and Agentic Workflows and API Actions and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Open library include?+

The full Open library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Answer parity across channels and Channel-appropriate formatting under Cross-Channel Agent Consistency); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Open or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Open areas and set them up in a Corsac workspace, where you can run every test case against Open or your own agent with your own data.