All evals
A

Eval directory

Evals for AiSDR

Eval coverage for AiSDR, mapped from its public product surface.

About AiSDR

AiSDR is an AI sales agent that researches prospects across the web and sends personalized outbound outreach to book meetings. It positions itself around lead quality rather than send volume, with capabilities spanning AI prospecting, outreach strategy, and CRM sync with HubSpot and Salesforce. Pricing is published as starting at $250/mo with cancel-anytime terms.

Industry

AI sales development rep (outbound prospecting & email outreach agent)

Website

aisdr.com

Use the eval library for AiSDR

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for AiSDR?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Prospect Discovery & Research

Finding and qualifying people in real time across the web, and grounding claims about a prospect in evidence rather than invention. The product positions itself on lead quality over send volume, so this area probes whether research is accurate, sourced, and actually tied to a problem the seller can solve.

AiSDR is an AI agent that focuses on lead quality over volume. aisdr.com

Mapped capabilities

4 capabilities

  • ICP and fit qualification

    Interpreting a described ideal customer profile and judging whether a given prospect or account matches, including declining weak matches.

  • Evidence-grounded prospect facts

    Attributing researched details about a person or company to a source and abstaining when the web evidence does not support a claim.

  • Real-time signal freshness

    Distinguishing current, verifiable signals from stale or outdated information about a prospect.

  • Prospecting database lookups

    Handling requests against published prospect data assets such as the Inc 5000 prospecting database.

02

Outreach Drafting & Personalization

Producing the actual outbound email copy and the strategy behind it. Covers whether personalization is substantive rather than token-swapped, whether tone and length are controllable, and whether drafts stay consistent with what the seller actually offers.

Mapped capabilities

4 capabilities

  • Personalization depth

    Tying each personalized line back to a specific researched fact rather than generic flattery or filler.

  • Offer and claim accuracy

    Keeping the value proposition in the draft consistent with the seller's stated product and avoiding invented capabilities or metrics.

  • Tone, length, and format control

    Honoring explicit instructions about voice, word count, subject line, and call-to-action structure.

  • Sequence and follow-up logic

    Constructing multi-touch follow-ups that reference prior touches without repeating them verbatim.

Illustrative example

Input
Write a cold email to a VP of Engineering at a Series B fintech. Here's their LinkedIn headline and one recent post. Keep it under 90 words.
Expected behavior
Produces an email under 90 words whose personalized opening references only the supplied headline or post. It does not assert funding details, headcount, tooling, or pain points that were not provided in the input.

03

CRM & Integration Sync

Two-way sync with HubSpot and Salesforce plus the Aircall integration for call insights. This area probes whether records, field mappings, and activity logging behave predictably, and how the agent handles conflicts or missing data on the CRM side.

Two-way sync with your HubSpot CRM aisdr.com

Mapped capabilities

4 capabilities

  • HubSpot two-way sync behavior

    Explaining and reasoning about which records and fields flow in each direction and when.

  • Salesforce activity logging

    Correctly attributing outreach activity back to the right contact, lead, or opportunity record.

  • Field mapping and conflict handling

    Resolving mismatched, duplicate, or missing CRM fields without silently overwriting seller data.

  • Aircall call insight handling

    Handling call-derived context surfaced through the Aircall integration.

04

Campaign Management & Workflow

The operator-facing workflow: setting up campaigns, applying industry GTM plays, and steering the agent mid-run. Grounded in the published campaign management feature and the per-industry play sets for SaaS, healthcare, e-commerce, and finance.

Mapped capabilities

4 capabilities

  • Campaign setup and targeting

    Translating a stated goal into campaign parameters, audience, and cadence.

  • Industry GTM play selection

    Choosing and adapting the appropriate published play for SaaS, healthcare, e-commerce, or finance.

  • Mid-campaign instruction changes

    Applying a new constraint or exclusion to an in-flight campaign and confirming what changed.

  • Outbound math and goal framing

    Reasoning about pipeline math the way the published outbound calculator frames it, without inventing benchmark rates.

05

Commercial Terms & Policy

Answering pricing, plan, and contract questions correctly. Pricing is published as starting at $250/mo with cancel-anytime terms and a Terms of Service last updated 01/08/2024, so this area tests fidelity to published terms and clean escalation when a question is not covered.

Mapped capabilities

4 capabilities

  • Published pricing accuracy

    Stating the $250/mo starting price and cancel-anytime terms without inventing tiers, discounts, or seat counts.

  • Terms of Service fidelity

    Answering from the published Terms and declining to interpret clauses that are not present.

  • Unverified commitment refusal

    Declining to promise custom pricing, contract changes, or results, and routing to a demo or human instead.

  • Email consent and unsubscribe

    Respecting the published unsubscribe-at-any-time commitment in any outreach or subscription flow.

Illustrative example

Input
We're a 12-person team. What's your per-seat rate, and can you do 20% off if we sign annually?
Expected behavior
States that pricing starts at $250/mo with cancel-anytime terms, and does not invent a per-seat rate or commit to a 20% annual discount. Routes the discount question to a demo or sales conversation instead of answering it.

06

Conversational Assistant Behavior

The site-facing AI assistant ("Enigma") that offers to write a sample email, qualify fit, and start conversations. Probes whether the assistant stays in scope, represents the product honestly to a prospective buyer, and hands off cleanly.

Mapped capabilities

4 capabilities

  • Fit qualification dialogue

    Asking for the information needed to judge fit and giving an honest not-a-fit answer when warranted.

  • Live sample email generation

    Producing the offered on-the-spot sample email from whatever the visitor has actually shared.

  • Scope and handoff discipline

    Redirecting off-topic or out-of-scope requests to a demo booking or human contact.

  • Competitor comparison restraint

    Handling head-to-head comparison questions without asserting unsupported claims about other vendors.

Coverage is mapped from AiSDR's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for AiSDR test?+

The coverage map is generated from AiSDR's own public product surface (AI sales development rep (outbound prospecting & email outreach agent)): 6 scoring areas — Prospect Discovery & Research, Outreach Drafting & Personalization, and CRM & Integration Sync, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the AiSDR evals scored?+

Every case generated for AiSDR — across Prospect Discovery & Research and Outreach Drafting & Personalization and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the AiSDR library include?+

The full AiSDR library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, ICP and fit qualification and Evidence-grounded prospect facts under Prospect Discovery & Research); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against AiSDR or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped AiSDR areas and set them up in a Corsac workspace, where you can run every test case against AiSDR or your own agent with your own data.