All evals
D

Eval directory

Evals for Docket

Eval coverage for Docket, mapped from its public product surface.

About Docket

Docket is an AI marketing agent embedded on B2B websites that engages visitors in unscripted voice or text conversations, qualifies their intent, books meetings, and syncs context into the CRM. It draws on a "Sales Knowledge Lake" plus CRM history and live intent signals to answer product and pricing questions, including on documentation pages. It is sold as an all-inclusive subscription priced by monthly website traffic rather than seats or conversations.

Industry

AI marketing agent for B2B website conversion

Use the eval library for Docket

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Docket?

6 scoring areas · 19 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Grounded Conversational Answering

Unscripted dialogue that reasons over the Sales Knowledge Lake instead of following a scripted chatbot playbook, staying inside what the customer's own product content supports.

Voice AI books 7.4x more meetings than text on the same traffic www.docket.io

Mapped capabilities

4 capabilities

  • Answer fidelity to source content

    Product and capability answers trace to the customer's ingested material rather than model priors.

  • Refusal and deferral on unsupported questions

    Questions outside the knowledge lake get an honest gap acknowledgement and a route to a human, not invention.

  • Multi-turn coherence

    Follow-ups resolve against earlier turns without restarting the qualification flow.

  • Voice and text parity

    The same question yields consistent substance across voice, text, slide and video modes.

Illustrative example

Input
Does your agent support on-premise deployment inside our own VPC with no outbound traffic?
Expected behavior
Does not assert a deployment model the ingested content does not cover. Says plainly that it cannot confirm that from available product material and offers to route the question to the Docket team or book time with a human.

02

Pre-Conversation Buyer Context

The profile Docket assembles before it speaks, drawn from CRM history, prior visit data, and live intent signals, and how that context shapes the opening exchange.

engages every website visitor in a real conversation www.docket.io

Mapped capabilities

3 capabilities

  • Known-account recognition

    Returning or CRM-known visitors get a warm, informed opening rather than a cold intro.

  • Cold-visitor handling

    Anonymous traffic with no history still receives a coherent open without fabricated familiarity.

  • Intent signal use

    Live signals and prior visits inform framing without overstating what is known about the visitor.

03

Intent Qualification and AQL

Turning a conversation into an Agent Qualified Lead: judging whether demonstrated intent and readiness are substantive enough to hand to sales, and declining to inflate the ones that are not.

a lead qualified through a real, substantive AI-powered conversation, demonstrating genuine intent and clear readiness for sales www.docket.io

Mapped capabilities

3 capabilities

  • Qualification signal extraction

    Need, fit, timing and role are captured from natural dialogue rather than an interrogation script.

  • Browsing versus buying discrimination

    Casual research is not promoted to a qualified lead.

  • Qualification without friction

    Technical or product questions get answered before gating on qualification questions.

04

Meeting Booking and CRM Sync

The conversion action itself and the handoff after it: calendar booking and the full conversation context written back into Salesforce or HubSpot.

CRM sync (Salesforce or HubSpot) www.docket.io

Mapped capabilities

3 capabilities

  • Calendar booking accuracy

    Proposed and confirmed times match availability and the visitor's stated constraints.

  • Context completeness in CRM

    The synced record carries the substantive conversation, not just contact fields.

  • Booking failure recovery

    Calendar or sync errors surface a clear fallback instead of a silently dropped lead.

05

Docs and Pricing Page Product Expert

Deep technical answering embedded on documentation, pricing, and product pages for self-serve evaluators who will not book a demo to get one question answered.

Docket embeds an AI product expert directly on your documentation and pricing pages www.docket.io

Mapped capabilities

3 capabilities

  • Technical depth on documented behavior

    API and configuration questions answered from actual product docs at evaluation-level specificity.

  • Documentation gap handling

    Questions the docs do not answer are named as gaps rather than smoothed over.

  • Self-serve continuity

    The agent accelerates the evaluation path instead of interrupting it with a demo gate.

06

Commercial Accuracy on Docket's Own Offer

How the agent represents Docket's traffic-based, all-inclusive subscription: tier thresholds, what each plan includes, and where it must stop short of quoting.

Priced based on your website traffic, not per seat or per conversation. www.docket.io

Mapped capabilities

3 capabilities

  • Tier and threshold accuracy

    Growth, Scale and Enterprise traffic bands and starting prices are stated as published.

  • Inclusion boundaries

    Features exclusive to Scale are not attributed to Growth.

  • Custom-pricing deferral

    Enterprise and non-standard requests route to a human rather than producing an invented quote.

Illustrative example

Input
We get about 12,000 visitors a month. Does the plan we'd be on include ABM account-level targeting and the video avatar?
Expected behavior
Places the visitor in Growth based on traffic and states that ABM account-level targeting and the video avatar are Scale features, not included at Growth. Offers the Scale tier or a demo as the path rather than implying the features are available.

Coverage is mapped from Docket's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Docket test?+

The coverage map is generated from Docket's own public product surface (AI marketing agent for B2B website conversion): 6 scoring areas — Grounded Conversational Answering, Pre-Conversation Buyer Context, and Intent Qualification and AQL, and more — spanning 19 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Docket evals scored?+

Every case generated for Docket — across Grounded Conversational Answering and Pre-Conversation Buyer Context and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Docket library include?+

The full Docket library is built on request. The coverage map spans 6 areas and 19 capabilities (for example, Answer fidelity to source content and Refusal and deferral on unsupported questions under Grounded Conversational Answering); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Docket or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Docket areas and set them up in a Corsac workspace, where you can run every test case against Docket or your own agent with your own data.