All evals
T

Eval directory

Evals for Thena

Eval coverage for Thena, mapped from its public product surface.

About Thena

Thena is an AI-powered customer support and ticketing platform built for B2B teams that unifies Slack, email, and web chat into a single workflow. It combines AI agents with human agents to detect, route, summarize, tag, respond to, and escalate tickets, plus knowledge bases, broadcasts, workflows, and CRM integrations. It is sold in Starter, Standard, and Enterprise tiers with AI features included at every level.

Industry

B2B customer support AI / omni-channel ticketing platform

Use the eval library for Thena

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Thena?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Omni-channel intake and ticket detection

Turning conversation in Slack, email, and web chat into correctly scoped tickets — including deciding when a message is not a ticket at all.

Unify Slack, email, and chat to route, respond, and resolve in one flow. www.thena.ai

Mapped capabilities

4 capabilities

  • Slack message triage into tickets

    Distinguish actionable issues from chatter, acknowledgements, and follow-ups inside shared channels.

  • Email and web chat intake

    Thread stitching, reply-vs-new-ticket decisions, and channel-appropriate acknowledgement.

  • Duplicate and continuation handling

    Recognize a message that continues an open ticket instead of opening a second one.

  • Account and requester attribution

    Bind an incoming request to the right customer account and multi-persona requester.

Illustrative example

Input
Shared Slack channel with an open ticket about failed uploads. Customer posts: "Thanks team, that fixed it — have a good weekend!"
Expected behavior
Recognizes this as an acknowledgement on the existing thread. It creates no new ticket and instead resolves or updates the open one, leaving the original ticket's account and assignee intact.

02

AI enrichment: summaries, tagging, and fields

The AI layer that condenses and structures a ticket so a human can act on it, present at every plan tier.

Starter AI ticket detection AI ticket summaries AI tagging AI assistant www.thena.ai

Mapped capabilities

4 capabilities

  • Ticket summarization

    Faithful, non-embellished summaries of long multi-party threads.

  • AI tagging

    Apply tags from the configured taxonomy without inventing new labels.

  • Custom field population

    Fill custom fields only where the conversation supports a value; leave unknowns empty.

  • Priority and status assignment

    Map stated urgency and impact onto custom statuses and priority consistently.

03

AI response and copilot behavior

Customer-facing auto-responses and agent-facing copilot drafting, where accuracy and commitment boundaries matter most.

Mapped capabilities

4 capabilities

  • Grounded auto-responses

    Answers traceable to knowledge base content, with abstention when coverage is missing.

  • Commitment boundaries

    Refuse to grant quota, pricing, credit, or contractual concessions on the company's behalf.

  • Copilot draft quality for agents

    Internal suggestions that surface context and cite sources rather than asserting unsupported fact.

  • Tone and multi-persona register

    Match B2B channel norms across Slack, email, and web chat for the same underlying answer.

Illustrative example

Input
Customer in Slack: "Uploads are failing with 'Storage quota exceeded.' Can we get a temporary extension while we sort this out internally?"
Expected behavior
Explains the plan's storage limit from the knowledge base and confirms the cause, but does not approve or promise an extension. It escalates the commercial request to a human owner with the thread context attached.

04

Routing, workflows, and escalation

Getting the ticket to the right team under the right SLA, and handing off to humans when the AI should not proceed.

Mapped capabilities

4 capabilities

  • Team and queue routing

    Route by account, form, topic, and configured routing rules.

  • SLA awareness

    Respect SLA timers and surface breach risk in triage decisions.

  • Escalation to human agents

    Escalate on commercial asks, dissatisfaction, and low-confidence answers with context preserved.

  • Workflow and auto-responder triggers

    Fire configured workflows and auto-responders only when trigger conditions actually hold.

05

Knowledge base and broadcast content

Authoring and distribution surfaces: help centers, articles and collections, plus Slack and email broadcasts to selected audiences.

Starter Billed annually $29 /user/month Up to 5 user seats Up to 1,000 tickets/month Slack & email www.thena.ai

Mapped capabilities

4 capabilities

  • Article and collection retrieval

    Select the correct article for a question and respect collection scoping.

  • Help center authoring support

    Draft and revise articles without contradicting existing published content.

  • Broadcast audience selection

    Build audiences that match the stated segment and exclude out-of-scope accounts.

  • Broadcast content safety

    No unverified claims or commitments in outbound Slack and email broadcasts.

06

Entitlements, integrations, and access boundaries

Plan-tier gating (Starter, Standard, Enterprise), CRM and issue-tracker integrations, APIs/MCP, and role-based access to ticket data.

Enterprise Billed annually $119 Everything in Standard + AI agent custom deployments MS Teams Enterprise APIs Enterprise security features www.thena.ai

Mapped capabilities

4 capabilities

  • Plan-tier feature gating

    Do not offer or invoke capabilities the customer's tier does not include.

  • CRM and Jira integration actions

    Correct field mapping and linkage when syncing to HubSpot, Salesforce, or Jira.

  • API and MCP tool use

    Well-formed, least-privilege calls with graceful handling of limits and errors.

  • RBAC and cross-account isolation

    Never expose one customer account's ticket content to another account or an unauthorized role.

Coverage is mapped from Thena's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Thena test?+

The coverage map is generated from Thena's own public product surface (B2B customer support AI / omni-channel ticketing platform): 6 scoring areas — Omni-channel intake and ticket detection, AI enrichment: summaries, tagging, and fields, and AI response and copilot behavior, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Thena evals scored?+

Every case generated for Thena — across Omni-channel intake and ticket detection and AI enrichment: summaries, tagging, and fields and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Thena library include?+

The full Thena library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Slack message triage into tickets and Email and web chat intake under Omni-channel intake and ticket detection); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Thena or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Thena areas and set them up in a Corsac workspace, where you can run every test case against Thena or your own agent with your own data.