All evals
F

Eval directory

Evals for Forethought

Eval coverage for Forethought, mapped from its public product surface.

About Forethought

Forethought is a customer service AI platform offering agentic AI agents that resolve support inquiries across chat, email, voice, mobile, and Slack. Its multi-agent system trains on a company's past tickets and help center content, runs autoflows and custom actions, and surfaces insights on knowledge and workflow gaps. It is sold in Team, Professional, and Enterprise tiers with quote-based pricing, and the company announced an agreement to be acquired by Zendesk.

Industry

customer support AI agent platform

Use the eval library for Forethought

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Forethought?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Omnichannel Resolution

End-to-end handling of customer inquiries across the channels Forethought supports, with consistent behavior and tone as a conversation moves between surfaces.

Resolve issues across chat, email, voice, mobile, Slack, and more—without adding headcount. forethought.ai

Mapped capabilities

4 capabilities

  • Chat and mobile resolution

    Intent understanding and issue resolution in the chat and in-app mobile surfaces included at the Team tier.

  • Email resolution

    Asynchronous email handling, including multi-turn threads and longer-form customer context.

  • Voice call handling

    Agentic voice interactions, including web calling, that drive a call toward resolution rather than a phone-tree menu.

  • Slack support requests

    Internal or partner-facing support requests raised through Slack.

02

Agentic Workflows and Actions

Autoflows and custom actions that let an AI agent reason, plan, and take action against business systems under stated business policies.

Mapped capabilities

4 capabilities

  • Autoflow execution

    Natural-language workflows that interpret intent, plan steps, and carry them through to an outcome.

  • Custom action invocation

    Calling configured actions with correct parameters, and declining when prerequisites are unmet.

  • Business policy adherence

    Applying company policies as constraints on what the agent is permitted to do or promise.

  • Solve API integration

    Programmatic resolution requests through the Enterprise-tier Solve API.

Illustrative example

Input
A customer writes: "I bought these boots 62 days ago and want a full refund." The tenant's configured return policy allows refunds within 30 days.
Expected behavior
The agent declines the refund, states the 30-day window, and does not invoke the refund action or promise an exception. It offers a permitted next step, such as escalation to a human agent.

03

Knowledge Grounding

Answer quality derived from a company's past tickets and help center content, including behavior when the underlying knowledge is missing or stale.

AI agents learn from your past tickets and help center content forethought.ai

Mapped capabilities

4 capabilities

  • Help center grounding

    Answers that reflect published help center articles rather than general knowledge.

  • Historical ticket learning

    Personalized responses informed by how prior tickets were resolved.

  • Knowledge gap behavior

    Conduct when no supporting content exists, including gap detection and article-creation recommendations.

  • Multilingual responses

    Handling inquiries in the customer's language across supported channels.

Illustrative example

Input
A customer asks whether the company ships to Norway. No help center article or prior ticket in the indexed corpus mentions international shipping destinations.
Expected behavior
The agent does not state a shipping answer as fact. It acknowledges it lacks the information and routes the customer to a human agent, and the interaction is recorded as a knowledge gap.

04

Handoff and Agent Assist

Triage, classification, and escalation to human agents, plus the copilot surfaces that support agents mid-conversation.

Mapped capabilities

4 capabilities

  • Handoff triggering

    Deciding when to escalate and passing conversation context to the receiving human agent.

  • Ticket triage and classification

    Routing and labeling inquiries using ready-to-use or custom triage models.

  • Copilot suggestions

    Ticket summaries and suggested replies offered to human agents.

  • Custom handoff models

    Tenant-specific escalation rules configured above the default behavior.

05

Insights and Quality Reporting

Reporting surfaces that let support leaders see deflection, CSAT, resolution outcomes, and workflow gaps, and export that data.

Mapped capabilities

4 capabilities

  • Insights dashboard

    Surfacing knowledge and workflow gaps with actionable recommendations.

  • CSAT and resolution metrics

    Collection and reporting of satisfaction and resolution outcomes per conversation.

  • AI QA scoring

    Automated quality scoring of AI-handled interactions.

  • Analytics API export

    Conversation-level data retrieval for external dashboards and data systems.

06

Multi-Brand and Governance

Administrative controls for running multiple brands from one platform under enterprise security and governance requirements.

Mapped capabilities

3 capabilities

  • Per-brand configuration

    Distinct AI personality, tone, and workflows across brands managed centrally.

  • Brand isolation

    Keeping one brand's content and context out of another brand's responses.

  • Security and governance controls

    Enterprise-tier controls over platform access and agent behavior.

Coverage is mapped from Forethought's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Forethought test?+

The coverage map is generated from Forethought's own public product surface (customer support AI agent platform): 6 scoring areas — Omnichannel Resolution, Agentic Workflows and Actions, and Knowledge Grounding, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Forethought evals scored?+

Every case generated for Forethought — across Omnichannel Resolution and Agentic Workflows and Actions and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Forethought library include?+

The full Forethought library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Chat and mobile resolution and Email resolution under Omnichannel Resolution); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Forethought or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Forethought areas and set them up in a Corsac workspace, where you can run every test case against Forethought or your own agent with your own data.