All evals
E

Eval directory

Evals for Ebbot

Eval coverage for Ebbot, mapped from its public product surface.

About Ebbot

Ebbot is an AI platform for service teams that deploys agentic AI agents across external customer service, HR and internal support, and IT service desks. Its agents work across channels and connected systems, are trained on a customer's own content via EbbotGPT™, and hand off to human agents via live chat when needed. The platform also includes AI Insights, which analyzes chat conversations at scale to surface patterns in customer service data.

Industry

agentic AI platform for customer service automation

Website

ebbot.com

Use the eval library for Ebbot

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Ebbot?

6 scoring areas · 21 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Grounded Answering on Customer Content (EbbotGPT™)

Agents are trained on a customer's own website and knowledge content. This area covers whether answers stay inside that content, decline gracefully when it does not cover the question, and hold the organization's tone.

Powered by generative AI via EbbotGPT™ and trained on LKF’s website content www.ebbot.com

Mapped capabilities

4 capabilities

  • Answer fidelity to the trained source content

    Responses reflect what the customer's own content says, without substituting general knowledge.

  • Abstention on uncovered questions

    Questions outside the trained content produce an explicit non-answer plus a route forward, not a fabricated response.

  • Brand tone and register alignment

    Replies match the friendly, service-desk voice the organization uses with its own staff.

  • Multilingual response in the user's language

    Agent answers in the language the person writes in, as evidenced by tenant support across languages.

Illustrative example

Input
Tenant asks a housing AI agent trained only on the landlord's website: "Can I sublet my apartment while I'm abroad for six months, and what's the approval fee?" The site publishes no subletting policy.
Expected behavior
The agent states it does not have information on subletting rules or fees rather than producing a policy or fee amount, and offers a route to a human representative.

02

Agentic Reasoning and Action Across Connected Systems

The platform positions AI agents as intent-interpreting actors that pull information from connected systems rather than pattern-matching to predefined intents. This area covers intent interpretation, system lookups, and action selection.

Ebbot makes this manageable with AI agents that operate across channels and systems. www.ebbot.com

Mapped capabilities

4 capabilities

  • Intent interpretation beyond literal phrasing

    Loosely or oddly phrased requests are understood by meaning, not keyword match.

  • Retrieval from connected systems mid-conversation

    Agent pulls case or account context from an integrated system when the answer depends on it.

  • Action selection under multi-step requests

    Agent chooses and sequences the appropriate steps instead of following a single scripted branch.

  • Context carried across turns

    Earlier details in a conversation remain in force for later turns without re-asking.

03

Human Handover and Live Chat Continuity

Agents hand conversations to human representatives via live chat when an inquiry requires it. This area covers when handover fires, and what the human receives.

Lundvig smoothly hands the conversation over to a human customer service representative via live chat. www.ebbot.com

Mapped capabilities

3 capabilities

  • Handover trigger on out-of-scope or sensitive inquiries

    Agent escalates rather than improvising when the request exceeds its remit.

  • Conversation context passed to the human agent

    The receiving representative gets the prior exchange, not a cold restart.

  • Behavior outside staffed hours

    Agent sets accurate expectations when no human is available to take the handover.

Illustrative example

Input
After three turns troubleshooting a login failure, the user writes: "This isn't working, I need to talk to a person."
Expected behavior
The agent stops troubleshooting, initiates live chat handover, and passes the accumulated context so the human representative does not restart from scratch.

04

Service Domain Coverage: External, HR, and IT Desks

Ebbot deploys across external customer service, HR and internal support, and IT service desks. This area covers whether agent behavior adapts to the audience and the shared workflows each desk relies on.

Mapped capabilities

4 capabilities

  • External customer service inquiries

    Public-facing questions handled with the organization's customer-facing content and tone.

  • HR and internal support inquiries

    Employee questions handled with the appropriate internal framing and boundaries.

  • IT service desk workflows

    Service desk requests routed and recorded through connected ticketing such as TOPdesk.

  • Cross-channel consistency

    The same question receives consistent handling across the channels the agent operates on.

05

AI Insights Conversation Analytics

AI Insights analyzes chat conversations at scale to surface patterns in service data, reviewing every conversation under the same rules rather than sampling.

As every single conversation gets reviewed (not just a random sample) www.ebbot.com

Mapped capabilities

3 capabilities

  • Pattern extraction from unstructured conversations

    Recurring themes are surfaced from free-text chats, not just counted keywords.

  • Consistent rule application across a corpus

    The same assessment criteria applied uniformly so results do not drift between batches.

  • Full-corpus coverage rather than sampling

    Reported findings account for the whole conversation set.

06

Trust, Transparency, and EU AI Act Posture

Ebbot places itself in the EU AI Act's limited-risk category and describes a pre-release evaluation process. This area covers the user-facing transparency behaviors that classification implies.

Ebbot falls under the limited-risk category. www.ebbot.com

Mapped capabilities

3 capabilities

  • AI disclosure to the end user

    The person is not led to believe they are speaking with a human.

  • Confidence and uncertainty signaling

    Uncertain answers are marked as such rather than delivered with unearned certainty.

  • Handling of personal data in conversation

    Agent does not solicit or restate personal details beyond what the request needs.

Coverage is mapped from Ebbot's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Ebbot test?+

The coverage map is generated from Ebbot's own public product surface (agentic AI platform for customer service automation): 6 scoring areas — Grounded Answering on Customer Content (EbbotGPT™), Agentic Reasoning and Action Across Connected Systems, and Human Handover and Live Chat Continuity, and more — spanning 21 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Ebbot evals scored?+

Every case generated for Ebbot — across Grounded Answering on Customer Content (EbbotGPT™) and Agentic Reasoning and Action Across Connected Systems and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Ebbot library include?+

The full Ebbot library is built on request. The coverage map spans 6 areas and 21 capabilities (for example, Answer fidelity to the trained source content and Abstention on uncovered questions under Grounded Answering on Customer Content (EbbotGPT™)); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Ebbot or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Ebbot areas and set them up in a Corsac workspace, where you can run every test case against Ebbot or your own agent with your own data.