All evals
C

Eval directory

Evals for Cognigy

Eval coverage for Cognigy, mapped from its public product surface.

About Cognigy

NiCE Cognigy is an enterprise CX AI platform for building and operating AI Agents that handle customer service across voice, chat, and messaging channels. It also offers an Agent Copilot that supports human service agents in-channel, plus an AI Ops Center for monitoring the AI workforce. The product ships on a frequent release cadence and is marketed to large brands with analyst-recognized leadership in conversational AI.

Industry

enterprise conversational & voice AI agents for customer service

Use the eval library for Cognigy

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Cognigy?

6 scoring areas · 22 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Voice AI Agent Conversations

Phone and voice interactions handled end to end by AI Agents, including how they hold a natural spoken conversation at scale and route callers to the right outcome.

With NiCE Cognigy Voice AI Agents, deliver empathetic and effortless phone conversations that scale. www.cognigy.com

Mapped capabilities

4 capabilities

  • Spoken intent understanding and resolution

    Interpreting caller requests over voice and driving them to a completed outcome rather than a restatement.

  • Routing accuracy on voice

    Selecting the correct destination, skill, or flow for the caller's stated need.

  • Empathetic and channel-appropriate phrasing

    Tone and response length suited to a spoken exchange rather than a text transcript.

  • Telephony/contact center connector behavior

    Behavior across voice integrations such as the Genesys Audio Connector path.

02

Chat & Messaging AI Agents

Digital and messaging channel agents serving customers across multiple channels and languages, including social messaging surfaces.

Equip service agents with essential tools and knowledge to swiftly and confidently manage customer queries across channels. www.cognigy.com

Mapped capabilities

4 capabilities

  • Cross-channel consistency

    Same customer request yields consistent answers across the digital and messaging channels in scope.

  • Multilingual handling and real-time translation

    Responding in the customer's language and preserving meaning through AI translation.

  • Multi-turn context retention

    Carrying earlier turns of a chat forward instead of re-asking for details already given.

  • Transactional task completion

    Completing service actions such as rebooking, refund, or appointment steps within the conversation.

03

Agent Copilot for Human Agents

In-channel assistance for human service agents: surfacing knowledge, drafting responses, and supporting agents handling queries across channels.

PCI DSS v4.0.1 Compliance: Enabling Secure Payment Handling with Enterprise AI Agents www.cognigy.com

Mapped capabilities

3 capabilities

  • Knowledge surfacing in-channel

    Bringing the relevant answer or article to the agent for the live customer query.

  • Suggested response quality

    Drafts an agent can send with minimal edits and without unsupported claims.

  • Conversation summarization for the agent

    Condensing the interaction so far without dropping or inventing key facts.

04

Knowledge Grounding & Resolution Accuracy

Whether agent answers stay tied to connected knowledge sources via retrieval-augmented generation, and how they behave when the knowledge base does not cover the question.

Mapped capabilities

3 capabilities

  • Answer grounded in retrieved sources

    Claims in the response trace to retrieved knowledge rather than model priors.

  • Abstention on out-of-scope questions

    Declines to answer and offers a next step when no source covers the request.

  • Coverage across a broad intent set

    Correct handling across the wide intent inventory an enterprise deployment carries.

Illustrative example

Input
Customer asks the chat AI Agent: "What's your refund window for items bought on a corporate account?" This policy detail exists in no connected knowledge source.
Expected behavior
The agent states it does not have that information and offers a handover or follow-up, rather than asserting a specific refund window or inventing a policy.

05

Escalation, Handover & Failure Recovery

What happens when the AI Agent cannot or should not proceed: transferring to a human, preserving context, and recovering from tool or integration failures.

Mapped capabilities

4 capabilities

  • Handover trigger correctness

    Escalates when warranted and avoids escalating requests it can resolve.

  • Context transfer on handover

    The receiving human agent gets the conversation history and collected details intact.

  • Recovery from tool or endpoint failure

    Degrades gracefully when a REST endpoint, webhook, or backend call fails.

  • Repair after misunderstanding

    Re-asks or corrects course when the customer signals the agent got it wrong.

06

Enterprise Controls, Security & AI Ops

The operational and compliance plane around the AI workforce: payment-data handling, endpoint and webhook security, privacy-safe analytics, model updates, and monitoring and automated quality evaluation.

strengthens security for the MCP Server Endpoint with OAuth 2.0 support www.cognigy.com

Mapped capabilities

4 capabilities

  • PCI-scoped payment handling

    Sensitive cardholder data is masked and not retained where it would breach PCI DSS v4.0.1 scope.

  • Endpoint and integration authentication

    REST endpoint auth, webhook security, and MCP Server OAuth 2.0 behavior on unauthenticated or malformed calls.

  • Privacy-safe analytics and logging

    Analytics and log payloads exclude data that governance settings mark as restricted.

  • Model switching and release stability

    Behavior holds through one-click LLM updates within a provider and across the frequent release cadence.

Illustrative example

Input
In a PCI-scoped payment step, a caller reads a 16-digit card number and CVV to the voice AI Agent to settle an outstanding invoice.
Expected behavior
The agent completes the payment step while the card number and CVV are masked everywhere they surface, with only a masked reference retained on the interaction record.

Coverage is mapped from Cognigy's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Cognigy test?+

The coverage map is generated from Cognigy's own public product surface (enterprise conversational & voice AI agents for customer service): 6 scoring areas — Voice AI Agent Conversations, Chat & Messaging AI Agents, and Agent Copilot for Human Agents, and more — spanning 22 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Cognigy evals scored?+

Every case generated for Cognigy — across Voice AI Agent Conversations and Chat & Messaging AI Agents and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Cognigy library include?+

The full Cognigy library is built on request. The coverage map spans 6 areas and 22 capabilities (for example, Spoken intent understanding and resolution and Routing accuracy on voice under Voice AI Agent Conversations); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Cognigy or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Cognigy areas and set them up in a Corsac workspace, where you can run every test case against Cognigy or your own agent with your own data.