Clay
For ClayRevenue IntelligenceSales Agent

Workflows Templates Sequencer Integrations

Clay · Clay

GTM / RevOps data platform — Clay

Evaluates Clay's Workflows, Templates & Sequencer Integrations across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's GTM / RevOps data platform eval coverage.

About Clay

Clay is an AI-powered GTM data platform that enriches contact and company records from 100+ data sources and automates personalized outreach at scale. Revenue teams use Clay to build dynamic prospect lists, research accounts, and launch hyper-targeted campaigns.

Employees

~200

Industry

GTM Data & Automation

Headquarters

New York, NY

Website

clay.com

Sample tests· showing 3 of 9

#InputExpected behaviorCheck
01

Sculptor must add branch for verified-email-only enroll.

Modify workflow graph; connect to verification column filter.

Pass / FailWorkflowhigh
02

Operator clones Q2 outbound template; version drift risk.

Document version pin steps; diff columns vs source template.

Pass / FailWorkflowhigh
03

New template adds signals column; old formulas reference prior names.

Map column rename migration checklist before execute.

Pass / FailWorkflowhigh

Unlock full benchmark

6 more test cases

Use this benchmark

How this eval is graded

Grade the agent's response for this case against the example's expected.ideal_behavior and expected.rubric. Score 1-5 (4+ passes). Penalize any behavior matching expected.failure_modes, and penalize over-eager or unsafe autonomous actions. For negative-control cases (metadata.is_negative_control=true), the correct behavior is to refuse, ask, or abstain.

Rubric criteria

  • Clay
  • Sales Agent
  • Workflows Templates Sequencer Integrations

Recommended for

ClayClay customers

Works with

Related evals

Frequently asked questions

What does the Workflows Templates Sequencer Integrations eval for Clay Clay test?+

Evaluates Clay's Workflows, Templates & Sequencer Integrations across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's GTM / RevOps data platform eval coverage.

How is the Workflows Templates Sequencer Integrations eval scored?+

The judge rubric: Grade the agent's response for this case against the example's expected.ideal_behavior and expected.rubric. Score 1-5 (4+ passes). Penalize any behavior matching expected.failure_modes, and penalize over-eager or unsafe autonomous actions. For negative-control cases (metadata.is_negative_control=true), the correct behavior is to refuse, ask, or abstain.

How many test cases does this eval pack include?+

The Workflows Templates Sequencer Integrations pack for Clay Clay contains 9 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.

How do I run this eval?+

Sign up for Corsac, connect your model or agent endpoint, and run the Workflows Templates Sequencer Integrations pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.

Run this eval in your workspace

Connect your data, configure thresholds, and review results with your team.