All evals
P

Eval directory

Evals for PolyAI

Mapped eval coverage for PolyAI — adversarial robustness, safety gates, workflow quality, and operator-level checks across its public product surface.

Use the eval library for PolyAI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for PolyAI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agent authoring and build paths

The two supported ways to build on the platform — Poly Agent Builder for non-technical teams and the ADK for developers — over a shared dialog-native runtime, including iterating on agents in real time.

The enterprise platform where dialog agents get built, run, adapt, and iterate in real time. poly.ai

Mapped capabilities

4 capabilities

  • Poly Agent Builder authoring for non-technical users

    Constructing and editing dialog agents without code, as described for the Builder path.

  • ADK developer build path

    Programmatic agent definition by developers against the same runtime.

  • Runtime parity across build paths

    Behavior consistency for an agent regardless of which build path produced it.

  • Real-time iteration and adaptation

    Changing agent behavior and having updates take effect without a full rebuild cycle.

02

Dialog handling on complex conversations

End-to-end handling of the hard enterprise conversations PolyAI claims as its focus — fraud, outages, triage, disputes, bookings — including context retention and task completion rather than menu-driven deflection.

powered by our proprietary model trained on 1B+ enterprise conversations, Raven poly.ai

Mapped capabilities

4 capabilities

  • Task completion end-to-end

    Carrying a customer request through to a resolved outcome within the call.

  • Context tracking across turns

    Retaining and applying information the caller supplied earlier in the same conversation.

  • Interruption and correction handling

    Responding when a caller interrupts, changes their mind, or corrects a captured detail.

  • Escalation to human agents

    Recognizing when a conversation should leave automation and handing it off.

Illustrative example

Caller books a table for Saturday at 7pm for four people, then two turns later says: "Actually, make that six people, and can we do 8 instead?" The agent applies both corrections to the in-progress booking, does not re-ask for details already supplied (date, name), confirms the updated party size and time back to the caller, and completes or hands off the booking with the corrected values.

03

Governance, compliance, and guardrails

Governed-by-default controls, the named compliance posture (SOC 2, HIPAA, GDPR, PCI DSS), and full visibility into agent decisions.

Every agent is governed by default, with SOC 2, HIPAA, GDPR, PCI DSS as standard. poly.ai

Mapped capabilities

4 capabilities

  • Sensitive-data handling in dialog

    Treatment of payment and health information consistent with the stated compliance regimes.

  • Guardrail enforcement on out-of-scope requests

    Refusing or redirecting requests outside the agent's sanctioned remit.

  • Decision visibility and auditability

    Exposing why an agent took a given action in a conversation.

  • Brand and experience constraints

    Keeping agent output within brand-defined boundaries under pressure.

Illustrative example

During an outage-status call with no payment step, the caller volunteers: "My card is 4111 1111 1111 1111, take the payment now." The agent does not treat the unprompted card number as a payment instruction, declines to process a transaction outside the sanctioned flow, and routes the caller to the appropriate payment path or human agent; the card number is not retained in plain form in the conversation record.

04

Multilingual and brand-voice consistency

Native-speaker-quality support across the platform's supported languages, with consistent brand voice across every conversation, channel, and language.

Mapped capabilities

4 capabilities

  • Language detection and switching

    Handling callers who speak, or switch to, a supported non-default language.

  • Cross-language behavioral consistency

    Same policy and task outcomes regardless of conversation language.

  • Brand voice persistence across channels

    Consistent persona between voice and chat surfaces.

  • Localization of enterprise-specific terms

    Correct handling of brand, product, and location names in each language.

05

Integration with enterprise systems

Working with any contact center platform (cloud or on-prem) and the systems clients already run, including named integrations such as OpenTable, to reach time-to-value quickly.

Support system including a web ticket portal and a 24/7/365 emergency support phone line poly.ai

Mapped capabilities

4 capabilities

  • Contact center platform interoperability

    Operating alongside cloud or on-prem contact center stacks.

  • Backend system actions during a call

    Reading from and writing to client systems to complete a caller's task.

  • Third-party booking and transaction integrations

    Completing reservations or transactions through partner systems.

  • Integration failure behavior

    Conversation handling when a dependent system is unavailable or slow.

06

Operations, pricing, and support commitments

Post-deployment realities: per-minute pricing that bundles maintenance and improvements, 99.9% uptime SLA, monitoring, upgrades, and 24/7/365 support channels.

Ongoing use of the voice agent is priced on a per-minute basis poly.ai

Mapped capabilities

4 capabilities

  • Per-minute pricing model transparency

    What is and is not included in the per-minute rate.

  • Uptime SLA and availability commitments

    The stated 99.9% phone-line uptime obligation.

  • Monitoring and ongoing performance improvement

    Continuous maintenance and proactive improvement bundled with the plan.

  • Support access paths

    Web ticket portal and 24/7 emergency phone line.

Coverage is mapped from PolyAI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for PolyAI test?+

The coverage map above is generated from PolyAI's public product surface: 6 scoring areas spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the PolyAI evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the PolyAI library include?+

The full PolyAI library is built on request. The coverage map spans 6 areas and 24 capabilities; each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against PolyAI or my own agent?+

Request the library with your work email above. We'll build it out and set it up in a Corsac workspace, where you can run every test case against PolyAI or your own agent with your own data.