All evals
LA

Eval directory

Evals for Level AI

Eval coverage for Level AI, mapped from its public product surface.

About Level AI

Level AI is a full-stack AI platform for customer experience and contact center operations, spanning quality assurance, conversation intelligence, voice of customer, real-time agent assist, coaching, screen recording, and autonomous AI workers. It runs on the company's own compute and a CX-native model layer called Level AI Latitude, described as seven purpose-built models with orchestration, evaluations, and feedback loops. Product modules include iCSAT for automated satisfaction scoring, AgentGPT for knowledge-grounded real-time answers, and analytics over unstructured omnichannel customer data.

Industry

contact center / customer experience AI platform

Use the eval library for Level AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Level AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Automated Quality Assurance

Scoring every interaction against evaluation criteria without sampling, and producing QA output a reviewer can audit and dispute.

Mapped capabilities

4 capabilities

  • Full-coverage interaction scoring

    Applies scorecard criteria across all interactions rather than a sampled subset.

  • Evidence-backed scoring rationale

    Each score ties to specific transcript moments a reviewer can verify.

  • Scorecard criteria interpretation

    Handles compliance, soft-skill, and process criteria distinctly, including not-applicable cases.

  • Coaching handoff from QA findings

    Turns scored gaps into targeted, evidence-backed development items.

02

Real-Time Agent Assist

In-conversation guidance, summaries, and knowledge-grounded answers delivered to a live agent (Agent Assist, AgentGPT).

AgentGPT is trained on your organization’s data, allowing it to auto-generate accurate answers to customer questions thelevel.ai

Mapped capabilities

4 capabilities

  • Knowledge-grounded answer generation

    Answers drawn from the organization's own documentation, with sourcing.

  • Refusal and escalation on knowledge gaps

    Declines or escalates when the knowledge base does not cover the question.

  • Live conversation summarization

    AI-generated summaries usable for handoff and post-call documentation.

  • Feedback-driven improvement loop

    Incorporates agent feedback signals on prior answers.

Illustrative example

Input
Agent asks AgentGPT: "Can this customer return a final-sale jacket bought 45 days ago?" The connected knowledge base states final-sale items are non-returnable and the standard return window is 30 days.
Expected behavior
Answers that the return is not permitted, and grounds the answer in the two applicable knowledge base rules rather than general retail knowledge or an invented exception policy.

03

Voice of Customer and iCSAT

Inferring satisfaction, effort, emotion, and resolution from interactions rather than survey responses, and surfacing dissatisfaction drivers.

AI mines 100% of your customer interactions automatically to provide a complete view of customer satisfaction thelevel.ai

Mapped capabilities

4 capabilities

  • Satisfaction inference from interactions

    Scores emotion, effort, and resolution without a customer-filled survey.

  • Low-satisfaction driver and root-cause detection

    Identifies frustration points and the underlying cause, not just the sentiment.

  • Trend tracking across segments

    Compares satisfaction over time and across locations or groups.

  • Reasoning and evidence for each score

    Attaches supporting customer language to the reported score.

Illustrative example

Input
A support transcript where a billing error is not fixed and the agent promises a callback, and the customer closes with "okay, thanks so much, you've been great."
Expected behavior
Reports the interaction as unresolved and returns a satisfaction score below neutral, attributing it to the unresolved billing issue rather than reading the friendly sign-off as a satisfied outcome.

04

Conversation Intelligence and Analytics

Extracting structured analytics from unstructured omnichannel customer data and assembling them into reports.

automatically extracts a wide range of analytics from your unstructured omnichannel customer data thelevel.ai

Mapped capabilities

4 capabilities

  • Omnichannel data unification

    Consolidates interactions arriving from many upstream platforms.

  • Topic and intent extraction

    Identifies what was said and why across large volumes of interactions.

  • Report and chart construction

    Builds requested views through the chart builder over available fields.

  • Screen-recording signals in analysis

    Incorporates observed agent effort and friction alongside conversation content.

05

Model Harness: Guardrails and Governance

The production control layer described for routing, guardrails, validation, and governance across the CX-native model layer.

A control layer for routing, guardrails, and governance in production thelevel.ai

Mapped capabilities

4 capabilities

  • Input understanding and safety screening

    Classifies and screens inbound requests before downstream handling.

  • Guardrail enforcement on generated output

    Blocks or constrains responses that violate configured policy.

  • Intent clarification before acting

    Requests disambiguation instead of guessing on underspecified inputs.

  • Output validation and evaluation loops

    Checks generated results against expected form and quality before release.

06

AI Workers and Virtual Agents

Autonomous agents acting across workflows and conversational agents handling interactions end to end.

Mapped capabilities

4 capabilities

  • Multi-step workflow execution

    Carries an interaction through the sequence of actions it requires.

  • Handoff to a human agent

    Transfers with context when autonomous handling is not appropriate.

  • Boundary adherence during autonomous action

    Stays within the actions the workflow authorizes.

  • End-to-end conversational resolution

    Resolves the customer's stated need without human intervention where in scope.

Coverage is mapped from Level AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Level AI test?+

The coverage map is generated from Level AI's own public product surface (contact center / customer experience AI platform): 6 scoring areas — Automated Quality Assurance, Real-Time Agent Assist, and Voice of Customer and iCSAT, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Level AI evals scored?+

Every case generated for Level AI — across Automated Quality Assurance and Real-Time Agent Assist and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Level AI library include?+

The full Level AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Full-coverage interaction scoring and Evidence-backed scoring rationale under Automated Quality Assurance); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Level AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Level AI areas and set them up in a Corsac workspace, where you can run every test case against Level AI or your own agent with your own data.