All evals
S

Eval directory

Evals for Salespeak

Eval coverage for Salespeak, mapped from its public product surface.

About Salespeak

Salespeak is an AI agent platform that turns a B2B website into a real-time expert, answering questions from human visitors and from AI agents like ChatGPT, Claude, Perplexity and Gemini that crawl the site on buyers' behalf. It combines branded chat, voice and email engagement with a trained sales knowledge base, intent and journey signals, an agent analytics dashboard, and CRM sync to HubSpot and Salesforce. The stated goal is one verified source of truth so every telling of a company's story — human or machine — stays consistent.

Industry

AI inbound sales agent / agent interaction platform for B2B websites

Use the eval library for Salespeak

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Salespeak?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Verified Answer Fidelity

Whether visitor-facing answers are drawn from the single verified source of truth, attributed to a source page, and free of superseded or contradictory claims when the underlying content is updated.

on our own site, AI agents made 37,000 page fetches in the last 90 days salespeak.ai

Mapped capabilities

4 capabilities

  • Source-grounded responses with page attribution

    Answers cite the page they were drawn from, in the style of the site's 'source: /pricing' pattern.

  • Superseded claim suppression after an update

    Once a fact is fixed in the knowledge base, older phrasing from blog posts, decks, or prior pages stops surfacing.

  • Conflict handling across contradictory site content

    Behavior when a pricing page, blog post, and product page disagree about the same fact.

  • Deferral on unverified or out-of-scope questions

    Refusing to assert facts the verified source does not cover, rather than improvising.

Illustrative example

Input
A visitor asks: "When does the integration ship? I read on your blog it was coming next quarter."
Expected behavior
The agent answers from the current verified source, stating the integration shipped in May, and attributes the answer to the page it came from. It does not repeat the superseded "next quarter" claim or present both versions as equally valid.

02

AI Agent Traffic Surface

How the platform serves and accounts for non-human traffic: agents like ChatGPT, Claude, Perplexity, and Gemini fetching pages in real time because a buyer just asked about the company.

Mapped capabilities

4 capabilities

  • Agent-directed answers at the edge

    Agent Optimizer responses served to crawling agents rather than only to embedded human chat.

  • Agent identification and attribution

    Distinguishing agent fetches by source (ChatGPT, Claude, Perplexity, Gemini) from human visitor sessions.

  • AI query volume accounting against plan limits

    AI queries metered separately from visitor conversations per the published plan structure.

  • Machine-readable answer surface (NLWeb)

    Standards-based exposure of the verified answers to agent consumers, per the site's NLWeb claim.

Illustrative example

Input
A Perplexity crawler fetches the pricing page in real time while researching the company on a buyer's behalf.
Expected behavior
The edge serves a grounded pricing answer citing the pricing page rather than an unattributed summary, and the fetch is recorded in Agent Analytics as agent traffic attributed to Perplexity and counted against the plan's AI query allowance.

03

Multi-Modal Visitor Engagement

The branded chat, voice, and email experience embedded on the customer's site, including how a conversation moves a buyer forward without forms or delays.

150 visitor conversations + 10,000 AI queries + $4 per additional conversation salespeak.ai

Mapped capabilities

4 capabilities

  • Branded in-page chat experience

    Embedded conversational answering on the customer's own site.

  • Voice engagement

    Spoken interaction as an alternate modality within the same agent.

  • Email engagement and reply handling

    Conversations continued over email, including replies to the customer's campaigns.

  • Consistency of the answer across modalities

    The same verified fact rendered identically in chat, voice, and email.

04

Sales Knowledge Base Training

Getting the agent trained on product and GTM goals quickly, keeping it current, and surfacing where the source content is thin.

Train your AI agent on your product and GTM goals in hours, not days salespeak.ai

Mapped capabilities

4 capabilities

  • Initial training from site and product content

    Standing up a usable knowledge base in hours per the product page claim.

  • One-click training update propagation

    A corrected fact updates everywhere it is told.

  • Content coverage gap identification

    Flagging critical questions the existing content cannot answer.

  • GTM-goal-aligned campaign configuration

    Custom campaign creation tied to stated go-to-market goals.

05

Intent, Journey & Analytics

The signals and reporting that turn conversations into decisions: who the visitor is, what they intend, and what the agent analytics dashboard reports back to the team.

Mapped capabilities

4 capabilities

  • Buyer intent identification

    Classifying intent from conversation and journey behavior.

  • IP matching and personalization

    Firmographic resolution used to tailor the experience.

  • Popular questions and topic reporting

    Aggregate view of what buyers and agents actually ask.

  • Website improvement recommendations

    Key areas for improvement derived from observed gaps.

06

CRM Sync, Privacy & Plan Governance

Integration into the existing GTM stack alongside the data protection and entitlement rules the platform commits to publicly.

Mapped capabilities

4 capabilities

  • Two-way HubSpot and Salesforce sync

    Native sync of conversation-sourced contact and firmographic data.

  • GDPR data minimization and subject requests

    Processor-role handling of personal data collected only as the customer directs.

  • Access controls and data encryption

    SOC 2 and GDPR aligned controls stated on the trust surfaces.

  • Usage limits and overage behavior

    Included conversations and AI queries, and what happens past the plan's included volume.

Coverage is mapped from Salespeak's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Salespeak test?+

The coverage map is generated from Salespeak's own public product surface (AI inbound sales agent / agent interaction platform for B2B websites): 6 scoring areas — Verified Answer Fidelity, AI Agent Traffic Surface, and Multi-Modal Visitor Engagement, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Salespeak evals scored?+

Every case generated for Salespeak — across Verified Answer Fidelity and AI Agent Traffic Surface and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Salespeak library include?+

The full Salespeak library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Source-grounded responses with page attribution and Superseded claim suppression after an update under Verified Answer Fidelity); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Salespeak or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Salespeak areas and set them up in a Corsac workspace, where you can run every test case against Salespeak or your own agent with your own data.