All evals
Clay

Eval directory · Revenue Intelligence

Evals for Clay

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Clay AI products.

About Clay

Clay is an AI-powered GTM data platform that enriches contact and company records from 100+ data sources and automates personalized outreach at scale. Revenue teams use Clay to build dynamic prospect lists, research accounts, and launch hyper-targeted campaigns.

Employees

~200

Industry

GTM Data & Automation

Headquarters

New York, NY

Website

clay.com

Use the eval library for Clay

All 70 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Related in Revenue Intelligence

All evals →

More Revenue Intelligence eval libraries

Coverage map

What would you measure for Clay?

7 areas · 70 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Claygent Ai Research Agent Grounding

Evaluates Clay's Claygent AI Research Agent Grounding across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's GTM / RevOps data platform eval coverage.

Mapped capabilities

10 scenarios

  • In-table Claygent prompts
  • MCP context connectors

Public sample case

Input
Claygent column in Accounts table must return CEO full name with source URL in cell note.
Expected behavior
Write prompt template referencing only company domain column; require citation column output.
Check
Pass / fail check

02

Crm Sync Write Back Safety

Evaluates Clay's CRM Sync & Write-back Safety across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's GTM / RevOps data platform eval coverage.

Mapped capabilities

10 scenarios

  • Salesforce field mapping
  • Human-edited field protection

Public sample case

Input
CRM sync mapping UI links Enriched Title column to Lead.Title.
Expected behavior
Produce mapping table screenshot checklist; never map formula debug column.
Check
Pass / fail check

03

Email Finder Verification Pipeline

Evaluates Clay's Email Finder & Verification Pipeline across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's GTM / RevOps data platform eval coverage.

Mapped capabilities

10 scenarios

  • Email waterfall ordering
  • SMTP verification stage

Public sample case

Input
Work Email waterfall column lists two generic email finder providers then SMTP verify stage.
Expected behavior
Order finders before verification; stop when verified email populated.
Check
Pass / fail check

04

Gdpr Ccpa Tcpa Compliance Fields

Evaluates Clay's GDPR / CCPA / TCPA Compliance Fields across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's GTM / RevOps data platform eval coverage.

Mapped capabilities

10 scenarios

  • EU person handling
  • Opt-out and marketing consent

05

Person Company Entity Resolution

Evaluates Clay's Person & Company Entity Resolution across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's GTM / RevOps data platform eval coverage.

Mapped capabilities

10 scenarios

  • Domain to company linkage
  • Duplicate person detection

06

Waterfall Enrichment Provider Ordering

Evaluates Clay's Waterfall Enrichment & Provider Ordering across 11 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's GTM / RevOps data platform eval coverage.

Mapped capabilities

11 scenarios

  • Credit order & provider tier
  • Stop-on-hit semantics

07

Workflows Templates Sequencer Integrations

Evaluates Clay's Workflows, Templates & Sequencer Integrations across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's GTM / RevOps data platform eval coverage.

Mapped capabilities

9 scenarios

  • Sculptor natural-language workflows
  • Template version & clone

Frequently asked questions

What do the Corsac evals for Clay test?+

Each eval pack tests Clay's public product surface — including Claygent Ai Research Agent Grounding, Crm Sync Write Back Safety, and Email Finder Verification Pipeline — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Clay evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 70 Clay cases — from Waterfall Enrichment Provider Ordering (11 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Clay library.

How many test cases does the Clay library include?+

The Clay eval library includes 70 graded test cases across 7 eval packs, the largest being Waterfall Enrichment Provider Ordering with 11 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Clay or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 7 Clay packs — Claygent Ai Research Agent Grounding and Crm Sync Write Back Safety and the rest — against Clay or your own agent with your own data.