All evals
U

Eval directory

Evals for Unthread

Eval coverage for Unthread, mapped from its public product surface.

About Unthread

Unthread is an AI-powered helpdesk that turns conversations in Slack, Teams, email, web chat, and API into tracked tickets. An agentic AI triages requests, answers from a self-updating knowledge base, and fires multi-step automations across connected tools like Jira, Salesforce, and Workday. It also provides analytics on deflection, CSAT, SLAs, and recurring issue trends.

Industry

Slack-native AI helpdesk and ticketing platform

Use the eval library for Unthread

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Unthread?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Omni-channel Intake & Conversational Ticketing

Turning messages from every supported channel into tracked tickets without breaking the conversation, and keeping them visible in one inbox.

Mapped capabilities

4 capabilities

  • Slack-native ticket creation from channels, DMs, and threads

    Conversation becomes a ticket with correct requester, source channel, and thread linkage; replies stay in the thread.

  • Email-to-ticket with threading and attachments

    Reply chains collapse onto one ticket; attachments preserved and associated.

  • Web chat, Teams, and API intake

    Tickets created via non-Slack surfaces carry the same fields and appear in the unified inbox.

  • Unified inbox consolidation across sources

    Cross-channel tickets are deduplicated, ordered, and filterable without leaving Slack.

02

Agentic AI Triage & Grounded Answering

The AI agent reading each incoming ticket, classifying it, answering from the knowledge base, and knowing when to hand off to a human.

Confidence-scored with human fallback unthread.io

Mapped capabilities

4 capabilities

  • Category, priority, and team auto-assignment

    Classification matches the request's actual intent and urgency signals in the message.

  • Knowledge-base-grounded answers

    Replies are traceable to indexed source content rather than unsupported generation.

  • Confidence scoring with human fallback

    Low-confidence or out-of-scope requests escalate to a human instead of guessing.

  • In-thread response with preserved context

    Answers land in the originating thread carrying prior conversation and requester context.

Illustrative example

Input
Slack DM to the helpdesk: "What's our policy on expensing a home office monitor?" No expense policy exists in the connected knowledge base.
Expected behavior
The agent does not invent a policy. It states it lacks a grounded answer, creates a ticket, and routes it to a human owner with the original message and requester context attached.

03

Self-Learning Knowledge & Gap Analysis

Documentation that maintains itself from resolved conversations, and detection of where docs are missing or stale.

Mapped capabilities

4 capabilities

  • Article drafting from resolved tickets

    Solutions captured in threads become usable draft articles with accurate steps.

  • Knowledge base sync with connected doc tools

    Updates propagate between Unthread and sources such as Notion without divergence.

  • Gap analysis for missing or outdated docs

    Recurring unanswered issues surface as documentation suggestions.

  • Answer freshness after source updates

    Superseded content stops being used once the underlying doc changes.

04

Automations & Integration Actions

Authoring workflows in natural language, visual flows, or code, and executing them correctly against connected systems.

Mapped capabilities

4 capabilities

  • Natural-language workflow authoring

    A described workflow compiles into the intended trigger, steps, and order.

  • Multi-step flows with conditional branching

    Branches select the correct path and multi-step chains run to completion.

  • Two-way sync with connected tools

    Ticket state and linked records in Jira, Salesforce, Linear, ServiceNow, and similar stay consistent in both directions.

  • Custom functions and API calls

    User-supplied code and external API steps execute with correct inputs and surfaced errors.

Illustrative example

Input
"When an employee asks about PTO in Slack, check their balance in Workday, file the request, and confirm the dates back in the thread."
Expected behavior
The builder emits a workflow triggered by a Slack PTO message that reads the Workday balance before filing the request, then posts a confirmation with the requested dates into the originating thread.

05

Service-Level Operations & Access Control

SLA tracking, routing to the right responder, approval paths, and permission boundaries around what each user can see or do.

Mapped capabilities

4 capabilities

  • SLA/SLO tracking and breach alerting

    Timers reflect ticket activity and breaches alert the responsible party.

  • On-call rotations and workload distribution

    Assignment respects rotation, availability, and expertise rules.

  • Escalations and approval requests

    Escalation paths and approval gates trigger on the defined conditions before action.

  • SSO and custom RBAC permissions

    Role boundaries govern visibility and actions; directory identity resolves to the right user.

06

Analytics & Reporting

Measuring what the system deflected, how customers felt, and which issues keep recurring — and getting that in front of the team.

Mapped capabilities

4 capabilities

  • AI deflection rate measurement

    Deflected volume is attributed to knowledge base or automation resolution rather than human handling.

  • Recurring issue grouping and trend lines

    Similar tickets cluster correctly and trends reflect actual volume shifts.

  • CSAT and NPS survey capture

    Surveys fire at the right lifecycle point and results tie to the source ticket.

  • BI export and Slack-delivered summaries

    Raw data pipes to external BI tools; summaries and alerts post into Slack accurately.

Coverage is mapped from Unthread's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Unthread test?+

The coverage map is generated from Unthread's own public product surface (Slack-native AI helpdesk and ticketing platform): 6 scoring areas — Omni-channel Intake & Conversational Ticketing, Agentic AI Triage & Grounded Answering, and Self-Learning Knowledge & Gap Analysis, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Unthread evals scored?+

Every case generated for Unthread — across Omni-channel Intake & Conversational Ticketing and Agentic AI Triage & Grounded Answering and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Unthread library include?+

The full Unthread library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Slack-native ticket creation from channels, DMs, and threads and Email-to-ticket with threading and attachments under Omni-channel Intake & Conversational Ticketing); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Unthread or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Unthread areas and set them up in a Corsac workspace, where you can run every test case against Unthread or your own agent with your own data.