All evals
Intercom

Eval directory

Evals for Intercom

Eval coverage for Intercom, mapped from its public product surface.

About Intercom

Intercom is a customer service platform that pairs a full-featured helpdesk with a natively integrated AI Agent called Fin. Human agents work from an omnichannel shared inbox with ticketing, a Copilot AI assistant, outbound onboarding messages, and AI-powered reporting, all on the same customer record as Fin. It is sold per seat plus per Fin outcome, and Fin can also be run on top of an existing helpdesk such as Salesforce.

Industry

AI customer service / helpdesk platform

Use the eval library for Intercom

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Intercom?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Fin AI Agent resolution

Fin answering customer questions across channels, grounded in the customer's own content, with configurable voice and a clean exit to a human when it should not answer.

Mapped capabilities

4 capabilities

  • Grounded answers from help content

    Answers email, live chat, phone and other channels using the customer's help center and connected content; declines or defers when content does not cover the question.

  • Tone and answer-length control

    Respects configured tone of voice and answer length settings without changing the substance of the answer.

  • Taking action on external systems

    Invoking connected systems to complete a request rather than only describing the steps.

  • Handoff to human agents

    Routes to agents in the preferred channel with conversation context intact; recognizes when handoff is the correct outcome.

Illustrative example

Input
End customer in live chat: "We're on Essential. What does a Fin resolution cost on top, and can we turn on SSO for our team?"
Expected behavior
Fin states that Fin outcomes start from $0.99 and that SSO and identity management is an Expert plan feature not available on Essential, then offers the upgrade path or hands off to sales rather than quoting an unlisted price.

02

Agent workspace and Copilot

The omnichannel shared inbox where human agents work, including Copilot's in-line assistance and the permissions that bound which content and conversations it may use.

“Increase agent efficiency by 31% with Copilot” www.intercom.com

Mapped capabilities

4 capabilities

  • Copilot answer sourcing

    Draws on conversation history, internal content, and external sources; attributes which source an answer came from.

  • Troubleshooting and onboarding support

    Helping agents work through complex issues and internal training or onboarding material.

  • Team Inboxes and omnichannel routing

    Conversations, tickets, channels, and customer data surfaced in one workspace across teams.

  • Copilot permissions and oversight

    Admin limits on which content and conversation data Copilot can pull from, and visibility into how agents use it.

03

Ticketing and escalation

Moving from conversation to ticket without losing context, and keeping both the internal team and the customer current as work progresses.

Mapped capabilities

4 capabilities

  • Customer tickets from conversations

    One-click conversion of a conversation to a ticket with context preserved.

  • Back-office tickets

    Separate tickets that carry the context other internal teams need to collaborate.

  • Tracker tickets for shared issues

    A single ticket covering an issue affecting many users, with updates propagated to linked conversations.

  • Ticket forms and state updates

    Collecting information upfront via bot or in-product forms; ticket states driving customer updates in Messenger and email.

Illustrative example

Input
Agent in the inbox: "Fourteen customers reported the same failed export in the last hour. Turn this conversation into something I can track."
Expected behavior
The system proposes a tracker ticket covering the shared issue rather than a customer or back-office ticket, links the affected conversations, and notes that state changes push updates to those customers in the Messenger and over email.

04

Outbound and proactive support

In-context automated messaging that onboards users and heads off known issues before they reach the inbox.

Mapped capabilities

4 capabilities

  • Onboarding surfaces

    Product Tours, Checklists, Mobile Carousels, and Tooltips delivered in-product without code.

  • Targeted journeys with Series

    Omnichannel message sequences built for defined customer segments in the visual builder.

  • Proactive issue notification

    Informing affected customers when an issue arises to reduce inbound volume at the source.

  • Audience targeting correctness

    Messages reaching the intended segment and honoring channel selection.

05

Reporting and AI insights

Measuring the combined AI and human operation: pre-built and custom reports, automatic conversation scoring, and the recommendations that follow from them.

“Make data-driven decisions faster with 12 pre-built reports” www.intercom.com

Mapped capabilities

4 capabilities

  • Pre-built report templates

    Holistic support overview, teammate performance, conversation topics, and the other curated templates.

  • Custom reports and filtering

    Chart configuration, advanced filters, drill-down, and time-period comparison.

  • Automatic conversation scoring and alerts

    Scoring and organizing every conversation across Fin and human agents; surfacing emerging trends and alerts.

  • Recommendations to close gaps

    Flagging missing content or data integrations that limit resolution, with one-click action.

06

Workflows, knowledge, and administration

The configuration layer that determines what Fin knows, who sees what, and which capabilities a workspace is entitled to.

“SSO & identity management HIPAA support Service level agreements (SLAs)” www.intercom.com

Mapped capabilities

4 capabilities

  • Workflows automation builder

    Automation rules and round robin assignment across team inboxes.

  • Help center and knowledge management

    Public, private, multilingual, and multibrand help centers as Fin's and Copilot's content base.

  • Security and identity controls

    SSO and identity management, HIPAA support, SLAs, and the privacy boundaries documented in the help center.

  • Plan entitlement and billing model

    Essential, Advanced, and Expert seat tiers, Lite seats, per-Fin-outcome charges, and pay-as-you-go channels.

Coverage is mapped from Intercom's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Intercom test?+

The coverage map is generated from Intercom's own public product surface (AI customer service / helpdesk platform): 6 scoring areas — Fin AI Agent resolution, Agent workspace and Copilot, and Ticketing and escalation, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Intercom evals scored?+

Every case generated for Intercom — across Fin AI Agent resolution and Agent workspace and Copilot and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Intercom library include?+

The full Intercom library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Grounded answers from help content and Tone and answer-length control under Fin AI Agent resolution); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Intercom or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Intercom areas and set them up in a Corsac workspace, where you can run every test case against Intercom or your own agent with your own data.