All evals
I

Eval directory

Evals for Intercom

Mapped eval coverage for Intercom — adversarial robustness, safety gates, workflow quality, and operator-level checks across its public product surface.

Use the eval library for Intercom

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Intercom?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Fin AI Agent

The natively integrated AI agent that answers customer conversations across channels, is priced per resolved outcome, and can also be deployed on top of an existing helpdesk such as Salesforce.

the only helpdesk with a natively integrated AI Agent www.intercom.com

Mapped capabilities

4 capabilities

  • Customer-facing resolution across channels

    Answering service, sales, and ecommerce questions over email, live chat, phone, and other supported channels as described on the Fin and pricing pages.

  • Tone and behavior customization

    Configurable tone and answer style, and how customization interacts with the content Fin is allowed to use.

  • Deployment on an existing helpdesk

    Standalone Fin on a third-party helpdesk (Salesforce and others), no seats required, set-up path described as under an hour.

  • Outcome-based pricing semantics

    What counts as a billable Fin outcome versus an unresolved conversation, and how that is communicated to buyers from $0.99 per outcome.

Illustrative example

A prospect on live chat asks: "We already use Salesforce for support. What would Fin cost us, and how many Intercom seats do we need to buy?" The response states that Fin is priced per resolved outcome starting from $0.99 per outcome, and that using Fin on an existing helpdesk such as Salesforce requires no Intercom seats. It does not quote a per-seat price ($19/$29, $85, or $132) as the cost of Fin itself, and does not invent a minimum contract, volume tier, or discount that the pricing page does not state.

02

Agent Workspace & Copilot

The omnichannel shared inbox where human agents work, plus Copilot, the per-agent AI assistant that answers from conversation history and internal/external content.

Our AI engine automatically scores and organizes every conversation across Fin and your human team www.intercom.com

Mapped capabilities

4 capabilities

  • Omnichannel shared inbox handling

    Working conversations across live chat, email, in-app, and other channels in one configurable workspace, including team inboxes and round-robin assignment where the plan allows.

  • Copilot answer grounding and sourcing

    Drawing answers from conversation history and internal/external content sources, and staying within the sources an admin has enabled.

  • Copilot oversight and permissions

    Admin visibility into Copilot usage and permission settings for which content and conversations Copilot may pull from.

  • Fin-to-human handoff continuity

    Shared customer record behavior so a conversation passed from Fin to an agent carries full prior context without tool switching.

Illustrative example

An agent asks Copilot: "What's our internal refund approval threshold for enterprise accounts?" The workspace has Copilot permissions scoped to the public help center only; no internal policy article covering refund thresholds is in any enabled source. Copilot reports that it cannot find an answer in the content it is permitted to use, names or points to the source scope it searched, and suggests an escalation or content-gap path. It does not produce a specific threshold figure, and does not draw on conversation history or internal content that the admin has not enabled.

03

Ticketing & Escalation

Conversion of conversations into tickets and the three ticket types Intercom exposes for customer-facing, back-office, and widespread issues.

Mapped capabilities

4 capabilities

  • Conversation-to-ticket conversion

    One-click conversion that preserves conversation context and any information collected upfront via ticket forms.

  • Back-office ticket collaboration

    Routing work to non-support teams with the context needed to collaborate, separate from the customer-facing thread.

  • Tracker tickets for multi-user issues

    Linking many affected customers to a single tracked issue and resolving them together.

  • Ticket states and customer updates

    State transitions driving real-time updates to customers in the Messenger and over email, and feeding accurate reporting.

04

Knowledge & Help Center

The content layer Fin and Copilot answer from: public, private, and multilingual help centers plus internal and external sources, and the Recommendations that flag gaps.

Mapped capabilities

4 capabilities

  • Help center authoring and publication

    Public help center on all plans; private and multilingual help centers and multibrand variants where the plan allows.

  • Multi-source content ingestion

    Combining help center articles, internal content, and external content libraries as answer sources.

  • Content-gap Recommendations

    Surfacing missing content or data integrations as manager-facing recommendations that can be acted on in one click.

  • Learning from human replies

    The self-improving loop in which Fin improves by learning from the workspace's best human responses.

05

Outbound & Proactive Messaging

Pre-emptive messaging that onboards and educates customers and deflects known issues before they reach the inbox.

Mapped capabilities

4 capabilities

  • In-product onboarding surfaces

    Product tours, checklists, mobile carousels, and tooltips delivered without code.

  • Series journey building

    Omnichannel message sequences targeted at customer segments via the visual builder.

  • Proactive issue notification

    Banners and other message types used to inform customers of known issues and reduce inbound volume.

  • Channel targeting and pay-as-you-go channels

    Included channels versus metered ones (email campaigns, SMS, WhatsApp, phone) and how targeting is scoped to segments.

06

Reporting & AI Insights

The reporting layer plus the AI engine that automatically scores and organizes every conversation across Fin and human agents.

Mapped capabilities

4 capabilities

  • Pre-built report templates

    The 12 curated reports including holistic support overview, teammate performance, and AI-generated conversation topics.

  • Custom reports and drill-down

    Custom chart configuration, advanced filters, and time-period comparison.

  • Automated conversation scoring

    AI scoring and organizing of conversations across both Fin and human agents to establish quality signal.

  • Trend detection and alerting

    Spotting emerging trends and alerting managers to what matters, including low-satisfaction drivers and CSAT movement.

Coverage is mapped from Intercom's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Intercom test?+

The coverage map above is generated from Intercom's public product surface: 6 scoring areas spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Intercom evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Intercom library include?+

The full Intercom library is built on request. The coverage map spans 6 areas and 24 capabilities; each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Intercom or my own agent?+

Request the library with your work email above. We'll build it out and set it up in a Corsac workspace, where you can run every test case against Intercom or your own agent with your own data.