All evals
T

Eval directory

Evals for Tidio

Eval coverage for Tidio, mapped from its public product surface.

About Tidio

Tidio is a customer service platform that combines a help desk, live chat, and the Lyro AI agent so automated and human support run in one place. Lyro is trained on a business's own verified data sources to answer customer questions in the brand's voice, while Flows handle proactive lead capture and sales automation. It is sold as a bundled AI + help desk plan or as a standalone AI agent that plugs into an existing help desk, priced by monthly conversation volume.

Industry

AI customer service and help desk software

Use the eval library for Tidio

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Tidio?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Lyro AI Agent Answer Grounding

Whether the AI agent answers from the business's own verified data sources, stays in the brand's voice, and declines or escalates rather than inventing an answer when coverage is missing.

Train Lyro AI Agent on verified data sources to make sure all his responses are reliable and helpful. www.tidio.com

Mapped capabilities

4 capabilities

  • Answering from verified sources

    Responses trace to the connected business data rather than general world knowledge.

  • Refusal and escalation on gaps

    Out-of-coverage questions produce an honest non-answer plus a handoff path, not a plausible guess.

  • Brand voice consistency

    Tone and phrasing hold across question types without drifting into generic assistant register.

  • Multi-turn context retention

    Follow-up questions resolve against earlier turns in the same conversation.

Illustrative example

Input
Customer asks the AI agent: "What's your return window for sale items bought during a promotion?" No promotional return policy exists in the connected knowledge sources.
Expected behavior
The agent states it cannot confirm the promotional return window from available information and offers to connect the customer to a human agent. It does not state a specific number of days or improvise conditions.

02

Help Desk & Ticketing Operations

The ticket lifecycle described on the Help Desk page: converting conversations into tickets, filtering, prioritization, and automated routing across channels from one place.

Brands that switch to Tidio Help Desk see an average 24% uplift in CSAT. www.tidio.com

Mapped capabilities

4 capabilities

  • Conversation-to-ticket conversion

    A chat or email thread becomes a ticket with the right context carried over.

  • Automated routing

    Incoming requests reach the correct queue or owner by rule.

  • Priority handling

    Priority is set and respected in ordering and filtering.

  • Email triage and reply

    Repetitive email requests are scanned, routed, and resolved without manual sorting.

03

Live Chat & Human–AI Handoff

How automated and human support share one conversation, including when the AI hands off, what the agent receives, and how coverage works outside staffed hours.

Mapped capabilities

4 capabilities

  • Handoff trigger correctness

    Escalation fires on complexity, frustration, or explicit request for a human.

  • Context transfer to agent

    The human picks up with prior turns and identified intent intact.

  • Offline and after-hours behavior

    Unstaffed conversations are captured rather than dropped.

  • Unified conversation view

    Automated and human turns remain a single readable thread.

04

Flows & Proactive Automation

Proactive, pre-conversation automation for lead capture, sales, and booking, including Lyro Smart Actions that take an action rather than only replying.

Mapped capabilities

4 capabilities

  • Lead capture flows

    Visitor details are collected and qualified before handoff.

  • Smart Actions execution

    An action is completed, not just described, when the customer asks for it.

  • Trigger targeting

    Proactive messages fire on the intended visitor condition and not indiscriminately.

  • Flow-to-conversation transition

    An automated flow yields cleanly to Lyro or a human when it runs out of scope.

05

Plans, Pricing & Usage Metering

The three separately metered counters on the pricing page — billable conversations, Lyro AI conversations, and visitors reached with Flows — plus the choice between the bundled platform plan and the standalone AI agent.

Mapped capabilities

4 capabilities

  • Counter distinction

    Billable conversations, AI conversations, and Flows reach are not conflated.

  • Bundled vs standalone fit

    Recommends the right mode for teams that already own a help desk.

  • Volume-tier reasoning

    Maps a stated monthly volume onto the corresponding tier without inventing prices.

  • Overage and growth questions

    Explains what happens past the top listed tier.

Illustrative example

Input
Buyer asks: "We handle about 800 support conversations a month and want AI to take roughly 300 of them. What do we need, and we already use Zendesk?"
Expected behavior
The answer treats billable conversations and Lyro AI conversations as two separate metered counters at those volumes, and flags the standalone Lyro AI Agent as the fit for an existing Zendesk help desk rather than the bundled platform plan.

06

Analytics, CSAT & Improvement Guidance

Unified reporting across AI and human conversations, plus the actionable suggestions Tidio surfaces for raising Lyro's CSAT and effectiveness.

The highest resolution rate on the market www.tidio.com

Mapped capabilities

4 capabilities

  • Unified AI + human reporting

    Metrics span both automated and agent-handled conversations.

  • Resolution rate interpretation

    Automation-rate figures are explained against the user's own data, not asserted as guaranteed.

  • CSAT improvement suggestions

    Recommendations point at specific, fixable gaps in training data or flows.

  • Coverage gap surfacing

    Questions the AI could not answer are made visible for follow-up.

Coverage is mapped from Tidio's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Tidio test?+

The coverage map is generated from Tidio's own public product surface (AI customer service and help desk software): 6 scoring areas — Lyro AI Agent Answer Grounding, Help Desk & Ticketing Operations, and Live Chat & Human–AI Handoff, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Tidio evals scored?+

Every case generated for Tidio — across Lyro AI Agent Answer Grounding and Help Desk & Ticketing Operations and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Tidio library include?+

The full Tidio library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Answering from verified sources and Refusal and escalation on gaps under Lyro AI Agent Answer Grounding); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Tidio or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Tidio areas and set them up in a Corsac workspace, where you can run every test case against Tidio or your own agent with your own data.