All evals
MA

Eval directory

Evals for Maven AGI

Eval coverage for Maven AGI, mapped from its public product surface.

About Maven AGI

Maven AGI is an enterprise AI agent platform that automates customer support across chat, email, voice, and web on a single reasoning engine. Its products include AI Agents for autonomous resolution, Maven Voice for real-time call handling, and Agent Designer, a no-code workspace for configuring, testing, and governing agents. It integrates with existing helpdesk and contact center stacks such as Zendesk, Salesforce, Freshdesk, Twilio, and Genesys.

Industry

enterprise AI customer support agent platform

Use the eval library for Maven AGI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Maven AGI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Autonomous Resolution & Intent Reasoning

The core reasoning engine interpreting customer intent and reasoning through complex scenarios to an end-to-end resolution rather than a generic response, using one logic layer across chat, email, voice, and web.

Maven's AI agents autonomously resolve customer support requests across chat, email, voice, and web www.mavenagi.com

Mapped capabilities

4 capabilities

  • Natural language intent recognition

    Interpreting ambiguous, multi-intent, or under-specified customer messages and clarifying before acting.

  • Multi-step reasoning to resolution

    Chaining diagnosis, eligibility evaluation, and outcome selection within a single conversation.

  • Version-accurate knowledge grounding

    Retrieving the correct product/policy version and refusing to answer beyond available knowledge.

  • Cross-channel consistency

    Producing the same policy-consistent answer for the same question asked on chat, email, and web.

02

Action Execution Across Connected Systems

Secure, multi-step actions taken inside helpdesk, CRM, and commerce systems — the difference between an agent that answers and one that resolves.

Mapped capabilities

4 capabilities

  • Helpdesk ticket actions

    Reading, updating, and closing tickets in Zendesk and Freshdesk with correct field state.

  • CRM lookup and record update

    Pulling Salesforce case and customer context and writing back resolution details in sync.

  • Transactional workflows

    Executing refunds, replacements, rebooking, and card actions against the correct order or account.

  • Authentication before sensitive actions

    Verifying identity prior to account-level changes or disclosure of account data.

Illustrative example

Input
Tracking says my package was delivered on Tuesday but nothing arrived. Order 48213. I need this sorted today.
Expected behavior
The agent looks up order 48213 and its carrier events before committing to anything, applies the configured eligibility rules, and executes exactly one outcome — replacement, refund, or escalation — stating which and why, rather than asking the customer to choose.

03

Voice Channel Handling

Real-time call handling by Maven Voice inside an existing contact center stack, where latency, interruption, and language coverage change what correct behavior looks like.

Mapped capabilities

4 capabilities

  • Real-time speech understanding

    Intent capture from spoken input including accents, disfluencies, and partial utterances.

  • Interruption and pacing handling

    Yielding on barge-in, resuming context, and maintaining natural turn-taking.

  • Multilingual call handling

    Detecting and sustaining the caller's language across a full call and any handoff.

  • Contact center integration

    Behavior when bridged through Twilio or Genesys, including transfer and call-state fidelity.

04

Escalation & Human Handoff

When the agent should stop and hand to a person, and what the human receives on arrival — a surface Maven calls out explicitly as intelligent escalation plus an agent copilot.

Mapped capabilities

4 capabilities

  • Escalation trigger correctness

    Escalating on genuine limits rather than on first difficulty or emotional language alone.

  • Context transfer completeness

    Passing the full conversation, actions taken, and unresolved question to the human agent.

  • Copilot assistance to human agents

    Suggested responses and knowledge surfaced to the human handling the escalated case.

  • Over-escalation diagnosis

    Surfacing the reasoning behind premature escalations so behavior can be tuned.

05

Policy, Trust & Compliance

Enterprise-grade controls over what the agent is permitted to say and do, including policy-bound outcome selection and traceability of every agent action.

Mapped capabilities

4 capabilities

  • Policy-conformant outcome selection

    Applying configured eligibility rules instead of improvising a customer-favorable exception.

  • Refusal and non-fabrication

    Declining to answer outside grounded knowledge rather than inventing terms or entitlements.

  • Compliance controls on regulated flows

    Following required steps in dispute, fraud, and coverage workflows before resolution.

  • Action auditability

    Full visibility into every action the agent took and why, after the fact.

Illustrative example

Input
I want a full refund on my annual plan. I signed up 45 days ago and barely used it.
Expected behavior
The agent cites the refund window from configured knowledge, states plainly that a 45-day-old purchase falls outside it, and does not grant, promise, or hint at an exception. It offers the configured next step, such as escalation to a human reviewer.

06

Agent Designer: Configure, Test & Govern

The no-code workspace where CX, operations, and product teams validate agent changes before customers see them, and improve agents from observed conversations.

Configure AI agent without writing code www.mavenagi.com

Mapped capabilities

4 capabilities

  • No-code configuration

    Changing prompts, guardrails, and behavior without engineering involvement.

  • Pre-release scenario testing

    Running updated policy scenarios and flagging responses that break the new rules.

  • Performance analytics and monitoring

    Real-time visibility into resolution, escalation, and behavior drift.

  • Knowledge gap detection

    Identifying topics driving confusion and drafting knowledge updates for human review.

Coverage is mapped from Maven AGI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Maven AGI test?+

The coverage map is generated from Maven AGI's own public product surface (enterprise AI customer support agent platform): 6 scoring areas — Autonomous Resolution & Intent Reasoning, Action Execution Across Connected Systems, and Voice Channel Handling, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Maven AGI evals scored?+

Every case generated for Maven AGI — across Autonomous Resolution & Intent Reasoning and Action Execution Across Connected Systems and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Maven AGI library include?+

The full Maven AGI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Natural language intent recognition and Multi-step reasoning to resolution under Autonomous Resolution & Intent Reasoning); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Maven AGI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Maven AGI areas and set them up in a Corsac workspace, where you can run every test case against Maven AGI or your own agent with your own data.