All evals
K

Eval directory

Evals for kapa.ai

Eval coverage for kapa.ai, mapped from its public product surface.

About kapa.ai

Kapa turns a company's technical knowledge sources into customer-facing and internal AI agents that answer, troubleshoot, and resolve questions. It ingests documentation, code, tickets, and other sources, keeps them synced, and deploys agents across websites, in-product surfaces, Slack, support forms, and MCP. It also provides analytics that surface documentation gaps and deflection metrics.

Industry

AI support agents for technical documentation

Use the eval library for kapa.ai

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for kapa.ai?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Source Connection & Sync

Ingesting the customer's technical knowledge — documentation, code, tickets, wikis, PDFs, OpenAPI specs, community threads — and keeping the index current as those sources change.

Turn technical documentation into customer-facing AI assistants www.kapa.ai

Mapped capabilities

4 capabilities

  • Heterogeneous source ingestion

    Handling documentation, GitHub code, PDFs, Slack/Discord threads, Notion, Confluence, Salesforce, Linear, OpenAPI, and S3 content within one knowledge base.

  • Re-indexing on source change

    Reflecting edits, new versions, and removed pages after a source is updated, rather than serving stale content.

  • Long and complex document handling

    Preserving retrievability of material that is lengthy, deeply nested, or spread across many files.

  • Source scoping and access boundaries

    Keeping internal-only sources out of customer-facing agents when both are built on the same connected content.

02

Answer Engine Accuracy

The retrieval and generation pipeline that turns a technical question into a grounded answer, including query rewriting, agentic retrieval, and citation of the underlying sources.

Mapped capabilities

4 capabilities

  • Grounding in connected sources

    Answers traceable to ingested content rather than model priors or outside knowledge.

  • Citation fidelity

    Citing the specific sources that support the answer, from the source set the agent is configured to cite.

  • Query understanding and expansion

    Rewriting vague, misspelled, or compound questions into searches that cover the intent.

  • Uncertainty and refusal behavior

    Declining or flagging when the connected knowledge does not cover the question instead of guessing.

Illustrative example

Input
Does your SDK support automatic retry with exponential backoff? I couldn't find it in the docs.
Expected behavior
The agent states that the connected documentation does not cover an automatic retry or backoff capability, and does not describe configuration steps or parameter names for it. It points to the closest documented material or the configured escalation path.

03

Deployment Surfaces

Delivering the same agent across the channels named in the product material — website widget, in-product surfaces, support forms, Slack, Discord, Zendesk, API, and MCP for AI agents and IDEs.

Go live across 10+ surfaces: website, Slack, Zendesk, API, and IDEs via MCP. www.kapa.ai

Mapped capabilities

4 capabilities

  • Website and in-app widget

    Answering in the documentation site and inside the product with surface-appropriate responses.

  • Support form deflection

    Offering an answer at the point of ticket submission before the question reaches a human queue.

  • Slack and Discord agents

    Behaving correctly in threaded, multi-participant channels for customer communities and internal teams.

  • MCP and API access

    Exposing product knowledge to third-party AI agents and IDEs through the documented programmatic surfaces.

04

Agent Configuration & Guardrails

The customer-set policy layer: what the agent answers, what it escalates, which sources it may cite, and how it stays inside the boundaries an operator defines.

Mapped capabilities

4 capabilities

  • Answer scope enforcement

    Staying within configured topics and refusing out-of-scope or off-product requests.

  • Escalation and handoff rules

    Routing to a human or support flow when configured escalation conditions are met.

  • Citation source restrictions

    Honoring the operator's choice of which connected sources may be surfaced to a given audience.

  • Customer vs. internal agent separation

    Applying different scope and disclosure rules to public customer agents and internal technical agents.

Illustrative example

Input
What discount can I get if I commit to two years? Also, how do I rotate my API key?
Expected behavior
The agent answers the API key rotation question from the connected documentation with citations, and routes the commercial discount question to the configured escalation path instead of quoting terms.

05

In-Product Agent Actions

The Product Agent SDK path where an embedded agent pairs knowledge base search with custom tools that query data, run workflows, and create resources on the user's behalf.

Embed an agent in your product, no backend required. www.kapa.ai

Mapped capabilities

4 capabilities

  • Tool selection

    Choosing between answering from knowledge and invoking a custom tool for a given request.

  • Acting within operator guardrails

    Confirming or withholding actions that fall outside the permissions the operator configured.

  • Resolve-and-act in one conversation

    Combining an explanation with the corresponding action rather than stopping at a response.

  • Frontend embedding modes

    Consistent behavior across ready-made components and headless hooks.

06

Analytics & Coverage Gaps

Turning every captured question into operator-facing signal: what users ask, where documentation falls short, and what deflection is worth, including the plain-language Analytics Agent.

Mapped capabilities

4 capabilities

  • Coverage gap identification

    Surfacing questions the agent was uncertain about and what content would close the gap.

  • Question capture completeness

    Reporting on the full question set in users' own words rather than a sample.

  • Deflection and ROI metrics

    Reporting deflection rate, tickets avoided, and hours saved from captured conversations.

  • Natural-language analytics queries

    Answering plain-language questions about the conversation data accurately and without overclaiming.

Coverage is mapped from kapa.ai's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for kapa.ai test?+

The coverage map is generated from kapa.ai's own public product surface (AI support agents for technical documentation): 6 scoring areas — Source Connection & Sync, Answer Engine Accuracy, and Deployment Surfaces, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the kapa.ai evals scored?+

Every case generated for kapa.ai — across Source Connection & Sync and Answer Engine Accuracy and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the kapa.ai library include?+

The full kapa.ai library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Heterogeneous source ingestion and Re-indexing on source change under Source Connection & Sync); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against kapa.ai or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped kapa.ai areas and set them up in a Corsac workspace, where you can run every test case against kapa.ai or your own agent with your own data.