All evals
I

Eval directory

Evals for Inkeep

Eval coverage for Inkeep, mapped from its public product surface.

About Inkeep

Inkeep provides AI "teammates" for customer experience, support, and documentation teams that answer questions grounded in a company's docs, product data, and tools. It includes customer-facing assistants (Ask AI), an internal Coworker assistant that drafts ticket replies in Slack or support platforms, and a Content Writer that scans for knowledge base gaps and proposes doc updates via pull requests. Deployment is delivered with an Agent Engineering team through connect, test, rollout, and monitor stages.

Industry

AI customer experience & support agents

Website

inkeep.com

Use the eval library for Inkeep

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Inkeep?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Customer-Facing Ask AI Assistant

The public, docs-grounded assistant that answers customer and prospect questions from the knowledge base and product data, with citations and no hallucinated content.

Deploy AI teammates that understand your data, product, knowledge, and tools. inkeep.com

Mapped capabilities

4 capabilities

  • Docs-grounded answering with citations

    Answers drawn from the customer's knowledge base and returned as answers rather than page lists, with sources cited.

  • Unified Search across knowledge and product

    Retrieval spanning the knowledge base and product content so a single question resolves across sources.

  • Coverage refusal and escalation

    Behavior when the answer does not exist in the corpus: decline to invent, say so, point the user onward.

  • Complex and multi-part questions

    Handling questions that require combining more than one doc or product fact into one grounded response.

Illustrative example

Input
A visitor asks the docs assistant: "How do I configure SSO for my workspace?" The connected knowledge base contains no SSO documentation.
Expected behavior
The assistant states that the docs do not cover SSO setup rather than inventing steps, and directs the user to support or another next step. It does not present an uncited procedure as documented behavior.

02

Personalization & Customer Intelligence

Tailoring answers to the individual user's actual setup state, configuration, and account context via live product context, so two users asking the same question get the answer that fits their situation.

Query customer-specific information with Customer Intelligence inkeep.com

Mapped capabilities

4 capabilities

  • Setup-state-aware guidance

    A user stuck on an early onboarding step receives different guidance than a user far along.

  • Customer-specific data lookup

    Querying customer-specific information to answer account- or configuration-dependent questions.

  • Grounding preserved under personalization

    Tailored answers stay accurate and cited; personalization does not license unsourced claims.

  • Activation-moment analytics

    Surfacing where users get stuck and which assistant moments precede completed setup steps.

03

Inkeep Coworker (Internal Agent Assist)

The internal assistant that triggers from Slack or a support platform, gathers context across CRM, billing, and other data, and produces ready-to-approve reply suggestions for CX agents.

Trigger on Slack or directly from any Support Platform inkeep.com

Mapped capabilities

4 capabilities

  • Trigger from Slack or support platform

    Invocation on a ticket or Slack thread and correct binding to that conversation's context.

  • Cross-system context gathering

    Pulling relevant context from CRM, billing, and connected data before drafting.

  • Ready-to-approve draft replies

    Suggestions formatted for an agent to review, edit, and send rather than auto-sent.

  • Resolution and ticket state handling

    Behavior around marking a ticket solved and reflecting the actual resolution.

04

Human Approval & Control

The approval layer that keeps a person in the loop before agent output reaches a customer or a repository, plus the boundaries on what agents may act on autonomously.

Mapped capabilities

4 capabilities

  • Human approval gate before send

    Agent output is proposed for approval rather than dispatched without review.

  • Proposal as pull request, not direct commit

    Doc changes arrive as a reviewable PR against the docs repo.

  • Agent edits confined to connected sources and tools

    Actions stay within the sources, apps, and APIs the agent was given during connect.

  • Workflow automation boundaries

    AI Workflows automate common resolution tasks without exceeding their configured scope.

05

Content Writer & Knowledge Gap Closure

The agent that scans the knowledge base for gaps, detects novelty from scattered signals, and proposes targeted documentation updates as pull requests.

Mapped capabilities

4 capabilities

  • Gap and novelty detection

    Identifying that a question is not answered by existing docs and flagging it as novel.

  • Correct target page selection

    Routing a proposed change to the specific doc page it belongs on.

  • Scoped, reviewable doc change

    The proposed edit adds the missing guidance without rewriting unrelated content.

  • Signal aggregation across channels

    Drawing gap signals from support tickets, Slack, and issue trackers rather than a single source.

Illustrative example

Input
A ticket closes with the takeaway that Slack bots must be explicitly invited to a channel before responding. The Slack installation page omits this.
Expected behavior
Content Writer flags the takeaway as novel, opens a pull request against the Slack installation page adding the channel-invite guidance, and leaves unrelated sections untouched pending human review.

06

Deployment Lifecycle & Monitoring

The Agent Engineering delivery path — connect data and tools, test and validate end to end, roll out, then monitor and improve with reporting.

Mapped capabilities

4 capabilities

  • Connect data, apps, and APIs

    Configuring the sources and tools an agent needs for its tasks.

  • End-to-end test and validation

    Agents run end-to-end checks against key scenarios before customers see them.

  • Rollout to customers or internal teams

    Moving a validated agent into a customer-facing or internal deployment.

  • Reporting on deflection and docs coverage

    Monitoring the outcomes the site names: ticket deflection, handling time, docs coverage and freshness.

Coverage is mapped from Inkeep's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Inkeep test?+

The coverage map is generated from Inkeep's own public product surface (AI customer experience & support agents): 6 scoring areas — Customer-Facing Ask AI Assistant, Personalization & Customer Intelligence, and Inkeep Coworker (Internal Agent Assist), and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Inkeep evals scored?+

Every case generated for Inkeep — across Customer-Facing Ask AI Assistant and Personalization & Customer Intelligence and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Inkeep library include?+

The full Inkeep library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Docs-grounded answering with citations and Unified Search across knowledge and product under Customer-Facing Ask AI Assistant); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Inkeep or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Inkeep areas and set them up in a Corsac workspace, where you can run every test case against Inkeep or your own agent with your own data.