All evals
C

Eval directory

Evals for Capacity

Eval coverage for Capacity, mapped from its public product surface.

About Capacity

Capacity is a unified CX automation platform that deploys AI agents across voice, chat, SMS, email and web to resolve customer and employee support requests. It pairs those agents with real-time agent assist, a self-updating knowledge base, helpdesk ticketing, post-interaction conversation intelligence and outbound campaigns. Pricing combines an annual platform fee across Core, Pro and Enterprise tiers with volume-based usage pricing for AI agents.

Industry

CX automation platform (AI customer support agents)

Use the eval library for Capacity

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Capacity?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Omnichannel AI Agent Resolution

Context-aware AI agents that resolve customer and employee requests across voice, chat, SMS, email and web, including specialist-to-specialist handoff and escalation to a human when the request exceeds what the agent should handle.

platform fee plus usage-based pricing for AI Agents, so you only pay for what you use capacity.com

Mapped capabilities

4 capabilities

  • Cross-channel request resolution

    Same intent handled consistently on voice, chat, SMS, email and web, with channel-appropriate formatting and length.

  • Specialist handoff within a call or thread

    A conversation that moves across topics (billing to delivery to refund) reaches the right specialist without restating context.

  • Escalation to a human

    Recognizing the hard or sensitive request and routing it out rather than continuing to attempt resolution.

  • Employee and IT Tier 0 self-service

    Shift-left handling of routine internal IT questions before Tier 2 or Tier 3 escalation.

02

Knowledge Orchestration and Grounding

The self-updating knowledge base that powers every agent and workflow: answers grounded in indexed customer content, new knowledge captured from resolved tickets, and mapping of responses to existing articles so the same answer is reused everywhere.

Copilot drafts responses, Autopilot automates replies and knowledge grows with every ticket. capacity.com

Mapped capabilities

4 capabilities

  • Grounded answering and abstention

    Answers stay inside indexed knowledge; gaps produce a handoff rather than an invented policy or figure.

  • Knowledge capture from resolved tickets

    A new answer written during ticket resolution is stored and reusable for the next similar inquiry.

  • Mapping to existing knowledge

    Recognizing that a response duplicates an existing article and associating it instead of forking a second answer.

  • Train-once consistency across surfaces

    The same question returns the same substantive answer whether asked of the virtual agent, assist or helpdesk.

Illustrative example

Input
Customer asks the chat agent: "What's your refund window for orders placed during a promotion?" No indexed article covers promotional-order refunds.
Expected behavior
The agent states it does not have that answer and offers a handoff to a human or a ticket, rather than producing a plausible refund window. Any adjacent general refund information it gives is clearly attributed to what is actually documented.

03

Real-Time Agent Assist

Live guidance delivered to human agents during an interaction: automatic customer context, sentiment analysis, conversation guidance and suggested next steps intended to improve speed, accuracy and consistency.

Mapped capabilities

4 capabilities

  • Customer context surfacing

    Relevant account and history context is presented at the right moment without the agent searching.

  • Sentiment analysis

    Detecting escalating frustration or satisfaction from the live transcript and reflecting it accurately.

  • Next-step and guidance suggestions

    Suggested actions are grounded in knowledge and applicable to the interaction in progress.

  • Suppression of unhelpful suggestions

    Withholding guidance when the transcript does not support a confident recommendation.

04

Helpdesk Ticketing and Automation

The ticketing platform: Copilot-drafted replies with selectable tone, Autopilot autonomous responses, Kanban tracking with custom swimlanes, SLA management and automated escalation and routing of incoming tickets.

Mapped capabilities

4 capabilities

  • Copilot draft quality and tone control

    Drafts follow the selected tone (concise, empathetic, enthusiastic, friendly, informative, professional) without changing the substance.

  • Copilot-to-Autopilot autonomy boundary

    Autopilot answers routine tickets automatically and holds back on tickets that warrant human approval.

  • Ticket routing and escalation

    Repetitive questions are deflected while issues that need a person are escalated to the right queue.

  • Kanban status and SLA tracking

    Ticket state, custom swimlanes and SLA timing are reported accurately.

05

Post-Interaction Intelligence and Quality

Conversation intelligence applied after the interaction: automatic scoring of interactions (AQA), CSAT-oriented insight, and the post-interaction workflows that turn what happened into improvement for the next interaction.

36.3B+ Automated interactions capacity.com

Mapped capabilities

4 capabilities

  • Automated interaction scoring

    Scores and rationales are traceable to what actually occurred in the transcript.

  • Conversation summarization

    Summaries preserve outcome, commitments made and unresolved items.

  • Trend and driver identification

    Recurring contact drivers are identified from interaction volume rather than asserted.

  • Post-interaction workflow triggering

    Follow-up actions fire on the conditions they were defined for, and not otherwise.

06

Commercial, Integration and Compliance Boundaries

What the platform will state about itself: Core, Pro and Enterprise entitlements, the platform-fee-plus-usage pricing model, 250+ out-of-the-box integrations and actions taken across connected systems, and the advanced compliance and HIPAA options positioned at Enterprise.

Mapped capabilities

4 capabilities

  • Tier entitlement accuracy

    Correctly attributing a capability to Core, Pro or Enterprise as published.

  • Pricing model explanation

    Describing platform fee plus volume-based AI Agent usage without inventing unpublished rates.

  • Cross-system actions via integrations

    Taking or declining an action in a connected system based on what the integration actually supports.

  • Compliance and data-handling posture

    Representing HIPAA and advanced compliance availability accurately, including as a packaged option.

Illustrative example

Input
Prospect asks: "We're planning to start on the Core plan. Does that include real-time Agent Assist for our live agents?"
Expected behavior
The response states that real-time Agent Assist is not in Core and is included starting with Pro, and describes Core as covering knowledge base, helpdesk and automations, AI Agent Builder, data indexing, integrations and analytics.

Coverage is mapped from Capacity's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Capacity test?+

The coverage map is generated from Capacity's own public product surface (CX automation platform (AI customer support agents)): 6 scoring areas — Omnichannel AI Agent Resolution, Knowledge Orchestration and Grounding, and Real-Time Agent Assist, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Capacity evals scored?+

Every case generated for Capacity — across Omnichannel AI Agent Resolution and Knowledge Orchestration and Grounding and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Capacity library include?+

The full Capacity library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Cross-channel request resolution and Specialist handoff within a call or thread under Omnichannel AI Agent Resolution); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Capacity or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Capacity areas and set them up in a Corsac workspace, where you can run every test case against Capacity or your own agent with your own data.