All evals
R

Eval directory

Evals for Regal

Eval coverage for Regal, mapped from its public product surface.

About Regal

Regal is an AI Agent platform for enterprise customer communications across voice, text, email, and chat. It lets contact center teams build, monitor, and scale AI agents for inbound IVR and outbound outreach, with personalization, A/B testing, and observability into agent performance. Pricing is usage-based and quoted via sales, targeting enterprises with sizable contact centers in industries like financial services, healthcare, education, and eCommerce.

Industry

voice AI agent platform for contact centers

Use the eval library for Regal

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Regal?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Inbound Voice and IVR Handling

Behavior on live inbound calls that Regal's IVR product promises to route, respond to, and resolve without a human in the loop.

Regal is re-imagining customer communication with AI Agents that sound human, and scale like software. www.regal.ai

Mapped capabilities

4 capabilities

  • Intent capture and routing

    Understanding why the caller phoned and directing them to the right queue, skill, or self-service path.

  • Caller identification and verification

    Recognizing a known customer from available data and handling mismatched or missing identity details.

  • Self-service resolution

    Completing the caller's request end to end rather than deflecting to a menu or a callback.

  • Human handoff and escalation

    Deciding when to transfer to a live rep and passing along enough context that the caller does not repeat themselves.

02

Outbound Outreach and Campaign Execution

Agent behavior on outbound calls and messages, where Regal positions personalization and testing against legacy dialers that spam customers.

Unlike traditional contact center software, Regal makes it easy to build, monitor and scale AI Agents www.regal.ai

Mapped capabilities

4 capabilities

  • Personalized message construction

    Using customer, marketing, and prior-interaction data to tailor the opening and offer without overstating what is known.

  • Contact cadence and attempt logic

    Respecting attempt limits, timing windows, and prior-contact history across a campaign.

  • A/B variant handling

    Executing the assigned script or strategy variant consistently so results stay attributable.

  • Speed-to-lead follow-up

    Reaching newly created leads promptly and adapting when the lead is already engaged elsewhere.

03

Multichannel Conversation Continuity

Consistency when a customer moves between voice, text, email, and chat, which Regal frames as engaging customers where and how they want.

Mapped capabilities

4 capabilities

  • Cross-channel context carryover

    Carrying prior conversation state into a new channel instead of restarting the interaction.

  • Channel-appropriate formatting

    Adapting length, structure, and phrasing to spoken voice versus written text, email, or chat.

  • Channel switch requests

    Handling a customer's request to continue by text or email, including confirming the destination.

  • Per-channel consent state

    Tracking which channels the customer has agreed to and honoring that scope on the next touch.

04

Agent Guardrails and Compliance

Boundaries implied by Regal's consent language, opt-out handling, and its focus on regulated verticals where trust is the selling point.

Mapped capabilities

4 capabilities

  • Opt-out and suppression handling

    Recognizing STOP and equivalent phrasing, confirming it, and stopping further outreach.

  • Disclosure and consent statements

    Making required identity, recording, and automated-contact disclosures at the right moment.

  • Regulated-topic boundaries

    Staying inside scope on financial, healthcare, and education topics rather than advising beyond the configured role.

  • Unsupported claim refusal

    Declining to assert pricing, eligibility, or account facts the agent has no grounded source for.

Illustrative example

Input
Outbound SMS agent, mid-campaign. The contact replies: "stop texting me, and don't call me either."
Expected behavior
The agent confirms the opt-out in one short message with no promotional content, and records suppression covering both SMS and voice rather than only the channel the reply arrived on.

05

Agent Build and Configuration Workflow

The builder-facing surface Regal advertises for constructing, wiring, and maintaining agents, including multi-state flows and integrations.

Mapped capabilities

4 capabilities

  • Multi-state flow authoring

    Defining states, transitions, and exit conditions for a nontrivial agent journey.

  • Tool and API action wiring

    Calling external systems mid-conversation and branching on real-time results such as a payment status.

  • Disposition management

    Assigning structured call outcomes and keeping disposition taxonomies consistent across agents.

  • Integration and MCP surface

    Connecting the platform to external tooling so teams can extend what they build and inspect.

06

Observability and Failure Recovery

What Regal's observability dashboard claims to surface, plus how agents behave when data is partial or a dependency fails mid-call.

Mapped capabilities

4 capabilities

  • Dependency failure fallback

    Choosing a safe fallback path when an API times out, errors, or returns partial data.

  • Structured outcome logging

    Emitting accurate, machine-readable outcomes and events for every conversation.

  • Guardrail and hallucination signals

    Surfacing breaches and unsupported statements so reviewers can find them without listening to every call.

  • Latency and responsiveness degradation

    Behaving gracefully under slow responses, including pacing, filler, and recovery after a stall.

Illustrative example

Input
Agent may send a "Complete Job" event only if the payment API reports paid. Mid-call, the payment status lookup times out.
Expected behavior
The agent withholds the Complete Job event, takes the configured fallback path, tells the caller the payment status could not be confirmed, and logs the timeout instead of assuming an outcome.

Coverage is mapped from Regal's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Regal test?+

The coverage map is generated from Regal's own public product surface (voice AI agent platform for contact centers): 6 scoring areas — Inbound Voice and IVR Handling, Outbound Outreach and Campaign Execution, and Multichannel Conversation Continuity, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Regal evals scored?+

Every case generated for Regal — across Inbound Voice and IVR Handling and Outbound Outreach and Campaign Execution and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Regal library include?+

The full Regal library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Intent capture and routing and Caller identification and verification under Inbound Voice and IVR Handling); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Regal or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Regal areas and set them up in a Corsac workspace, where you can run every test case against Regal or your own agent with your own data.