All evals
C

Eval directory

Evals for Coldreach

Eval coverage for Coldreach, mapped from its public product surface.

About Coldreach

Coldreach is an AI SDR platform that researches accounts in real time to surface buying signals matching a customer-defined ICP, then writes and sends personalized outbound email. Users define what counts as intent for their product, and the AI monitors accounts across multiple intent data sources to build targeted lead lists and sequences. The company is backed by Y Combinator.

Industry

AI SDR / outbound sales automation

Use the eval library for Coldreach

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Coldreach?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Custom Intent Signal Definition

Coldreach lets users define what counts as intent for their specific product and monitors accounts against that definition across 5+ intent data sources. This area covers whether a user-authored intent prompt is interpreted faithfully rather than loosened into generic buying interest.

Coldreach researches 113M+ accounts in real time to find only relevant leads. coldreach.ai

Mapped capabilities

4 capabilities

  • Custom intent prompt interpretation

    Translating a user's product-specific definition of intent into consistent account-level criteria without broadening it.

  • Multi-source signal reconciliation

    Combining evidence from several intent data sources, including cases where sources conflict or only one weakly fires.

  • Signal recency and timing judgment

    Distinguishing a live, current signal from a stale one when surfacing an account as ready for outreach.

  • Disqualifying and negative signal handling

    Respecting exclusions in the user's definition so accounts that match on surface keywords but fail a stated exclusion are dropped.

Illustrative example

Input
Intent prompt: qualify only accounts hiring a dedicated data engineer AND running Snowflake. Account evidence: an open data engineer role; no warehouse technology mentioned anywhere in the research bundle.
Expected behavior
The account is not qualified. The output states that the Snowflake condition was not met and that no warehouse evidence was found, rather than inferring it from the hiring signal.

02

Account Research and Lead Qualification

The core claim is that the AI SDR does real research on 113M+ accounts before deciding an account is worth contacting. This area covers whether qualification decisions are evidence-backed, attributable, and appropriately withheld when research is thin.

build targeted lead lists matching your criteria and start reaching out, with AI scanning 5+ data sources coldreach.ai

Mapped capabilities

4 capabilities

  • ICP fit assessment against stated criteria

    Judging an account against explicit firmographic and product-fit criteria supplied by the user.

  • Evidence citation for each qualification call

    Attaching the specific researched evidence that triggered a qualify or disqualify decision.

  • Contact selection within a qualified account

    Choosing which persona or role to reach given the signal that fired.

  • Abstention on insufficient evidence

    Declining to qualify or personalize when the research returned nothing substantive, rather than guessing.

03

Lead List and Campaign Construction

Coldreach markets building targeted lead lists in minutes with every lead vetted before outreach begins. This area covers the translation from stated criteria to an actual list and the hygiene of that list before a campaign starts.

Mapped capabilities

4 capabilities

  • Criteria-to-list translation

    Producing a list whose members all satisfy the criteria as written, including compound and negated conditions.

  • Duplicate and prior-contact suppression

    Avoiding repeat contacts of the same account or person across overlapping campaigns.

  • Precision versus volume tradeoff

    Behavior when strict criteria yield a small list, including whether criteria are silently relaxed to fill quota.

  • Source coverage transparency

    Reporting which data sources were scanned and which contributed to a given lead.

04

Personalized Message Generation

Every email is claimed to be personalized to the buying signal while keeping the customer's tone via custom instructions and avoiding a robotic or templated feel. This area covers the writing step itself.

AI SDR personalizes every email according to buying signals. coldreach.ai

Mapped capabilities

4 capabilities

  • Signal-anchored relevance

    Grounding the opener and ask in the specific signal that qualified the account, not generic flattery.

  • Custom tone instruction adherence

    Following user-supplied voice, length, and style constraints across generated messages.

  • Anti-template variation

    Producing materially different emails across accounts in the same campaign rather than one skeleton with swapped tokens.

  • Single clear ask construction

    Ending with one specific, easy-to-answer request appropriate to a cold first touch.

Illustrative example

Input
Research bundle contains exactly one signal: a public job posting for a compliance analyst. Tone instruction: plain, under 90 words, no exclamation marks. Write the first-touch email.
Expected behavior
The email opens on the compliance analyst hiring signal, introduces no other company facts, respects the tone and length limits, and closes with one specific question the recipient can answer in a sentence.

05

Sequence and Follow-Up Behavior

Coldreach runs sequences, and its published guidance frames follow-ups as adding new context rather than nudging. This area covers multi-touch behavior over time and the rules for stopping.

has a 3.8% human reply rate across 500,000+ AI SDR emails excluding auto-replies coldreach.ai

Mapped capabilities

4 capabilities

  • Cadence and step-count discipline

    Spacing and total touches in a sequence, including the interval before the first follow-up.

  • New information per touch

    Each follow-up adding a fresh observation, trigger, or proof point instead of a check-in restatement.

  • Reply and auto-reply handling

    Distinguishing a human reply from an auto-response and stopping the sequence appropriately.

  • Mid-sequence disqualification exit

    Halting outreach when new research invalidates the original signal or fit.

06

Factual Grounding of Outbound Claims

Because messages are sent to real prospects under the customer's name, fabricated account details or unsupported statistics are a direct brand risk. This area covers whether generated research summaries and emails stay inside what was actually retrieved.

Mapped capabilities

3 capabilities

  • No fabricated account facts

    Refraining from inventing funding events, headcount, tooling, or initiatives not present in the research bundle.

  • Metric and proof-point sourcing

    Attaching a source to any performance or comparison claim used in an email.

  • Restraint on sensitive personal details

    Keeping personalization to professional, business-relevant context.

Coverage is mapped from Coldreach's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Coldreach test?+

The coverage map is generated from Coldreach's own public product surface (AI SDR / outbound sales automation): 6 scoring areas — Custom Intent Signal Definition, Account Research and Lead Qualification, and Lead List and Campaign Construction, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Coldreach evals scored?+

Every case generated for Coldreach — across Custom Intent Signal Definition and Account Research and Lead Qualification and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Coldreach library include?+

The full Coldreach library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Custom intent prompt interpretation and Multi-source signal reconciliation under Custom Intent Signal Definition); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Coldreach or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Coldreach areas and set them up in a Corsac workspace, where you can run every test case against Coldreach or your own agent with your own data.