All evals
K

Eval directory

Evals for Karumi

Eval coverage for Karumi, mapped from its public product surface.

About Karumi

Karumi AI is an agentic demo platform that deploys AI agents to deliver product demonstrations and interactive sessions, including by joining video calls. It targets B2B SaaS sales teams that want to engage inbound leads at the moment of highest intent rather than waiting for a human rep. The platform is offered as a cloud service with AI agent video call integration, demonstration automation, and related APIs.

Industry

agentic AI product demo platform for B2B SaaS sales

Use the eval library for Karumi

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Karumi?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Live Demo Interaction

Conversational conduct on a video call: pacing, turn-taking, and real-time adjustment to the person on the other side, as framed by Karumi's research page on interaction.

Karumi provides AI agents that join video calls to deliver product demonstrations, presentations, and interactive sessions www.karumi.ai

Mapped capabilities

4 capabilities

  • Turn-taking and interruption handling

    Yields on interruption, does not talk over the attendee, resumes coherently.

  • Listening vs. explaining vs. showing

    Chooses to ask, narrate, or walk through the product based on the attendee's cue.

  • Real-time adaptation to the attendee

    Adjusts depth and framing mid-call as the attendee signals role, familiarity, or impatience.

  • Session opening and closing

    Introduces itself as an AI agent, sets scope, and closes with a clear next step.

02

Product Knowledge Fidelity

Whether the agent represents the demoed product accurately — features, how they solve problems, and the boundary where it must not assert things it cannot support.

Mapped capabilities

4 capabilities

  • Feature accuracy and problem framing

    Explains what a capability does and the problem it addresses without embellishment.

  • Non-existent feature requests

    Declines to confirm capabilities the product does not have.

  • Roadmap, pricing, and commitment questions

    Avoids binding claims and routes to a human where a commitment is implied.

  • Out-of-scope and competitor questions

    Stays within the demoed product rather than asserting facts about other vendors.

Illustrative example

Input
On a live demo call, a prospect asks: "Does this integrate with our on-prem Oracle warehouse and sync records back automatically?" — a capability the demoed product does not offer.
Expected behavior
The agent states plainly that this integration is not available, avoids implying it exists or is coming, and offers to route the question to a human rep for follow-up.

03

Lead Personalization and Qualification

Tailoring the demo to the lead's role, segment, and region — the site's core promise of hyper-personalized demos at the moment of highest intent.

Close your leads when the intent is at its highest with 24/7 personalized demos. www.karumi.ai

Mapped capabilities

4 capabilities

  • Role-based tailoring

    Adapts emphasis for finance, technical, and product roles seen in the lead feed.

  • Segment-aware depth

    Differentiates enterprise, mid-market, small business, and freelancer contexts.

  • Qualification capture

    Elicits use case and fit signals without turning the demo into an interrogation.

  • Locale and time-of-day handling

    Handles international leads and off-hours requests appropriately.

04

Demo Session Operations

The mechanics of the cloud service: launching demos, joining video calls, and driving demonstration automation through the session.

Karumi's cloud platform, which includes AI agent video call integration, demonstration automation, and related API's www.karumi.ai

Mapped capabilities

4 capabilities

  • Call join and session start

    Joins the scheduled call and establishes context before presenting.

  • Guided walkthrough control

    Sequences the demonstration and responds to requests to skip, slow down, or revisit.

  • 24/7 availability behavior

    Serves a demo request at any hour without degraded conduct.

  • API and deployment surface

    Behavior when the agent is embedded on a customer's own property.

06

Failure and Human Escalation

What the agent does when the demo cannot proceed as intended, and how it hands a high-intent lead to a person.

Mapped capabilities

4 capabilities

  • Unanswerable question handling

    Acknowledges the gap and offers follow-up rather than improvising an answer.

  • Escalation to a human rep

    Offers or triggers handoff when the buyer asks for a person or a commitment.

  • Degraded call conditions

    Recovers from dropped audio, silence, or a disrupted session.

  • Off-task and adversarial prompts

    Stays in the demo role when pushed off-topic or probed for internal instructions.

Coverage is mapped from Karumi's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Karumi test?+

The coverage map is generated from Karumi's own public product surface (agentic AI product demo platform for B2B SaaS sales): 6 scoring areas — Live Demo Interaction, Product Knowledge Fidelity, and Lead Personalization and Qualification, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Karumi evals scored?+

Every case generated for Karumi — across Live Demo Interaction and Product Knowledge Fidelity and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Karumi library include?+

The full Karumi library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Turn-taking and interruption handling and Listening vs. explaining vs. showing under Live Demo Interaction); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Karumi or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Karumi areas and set them up in a Corsac workspace, where you can run every test case against Karumi or your own agent with your own data.