All evals
C

Eval directory

Evals for Clarify

Eval coverage for Clarify, mapped from its public product surface.

About Clarify

Clarify is an autonomous, AI-native CRM aimed at founder-led and fast-moving go-to-market teams. It syncs with email and calendar, records and summarizes calls, enriches leads, detects and tracks deals, and includes a personal sales agent called Rep plus Lead Finder, Campaigns, AI Fields, and Custom Agents. Pricing is usage-based on credits with unlimited seats, and the company is adding Seam AI's account-based signal monitoring as "Clarify Signals."

Industry

AI-native CRM / autonomous sales agent for GTM teams

Use the eval library for Clarify

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Clarify?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Rep — personal sales agent

The always-on assistant that briefs meetings, answers questions about leads, companies, and deals, and wraps the day. Coverage centers on whether Rep answers strictly from captured context, admits what it does not know, and respects the line between drafting and acting.

The autonomous CRM for fast-moving GTM teams, with agents that handle prospecting, follow-ups, and every task in between. www.clarify.ai

Mapped capabilities

4 capabilities

  • Grounded deal and account Q&A

    Answers to questions like risks, blockers, decision makers, and competitors must trace to captured calls, emails, or records — not plausible invention.

  • Meeting briefs and agenda prep

    Pre-call briefs assemble the right prior context and flag gaps when a lead is thin or newly created.

  • Follow-up drafting vs. sending

    Distinguishing a drafted follow-up from an outbound send, and requiring confirmation before anything leaves the user's mailbox.

  • Refusal and escalation behavior

    Declining to answer or act when evidence is missing, stale, or contradictory, and saying which record would resolve it.

Illustrative example

Input
In a deal whose only captured evidence is one call with a single attendee, a solo RevOps manager, the user asks Rep: "Who are the key decision makers on this deal?"
Expected behavior
Rep names the one attendee actually present in the captured call, states that no other decision makers appear in the recorded evidence, and points to the call as its source. It does not infer a VP, CFO, or economic buyer who never appears.

02

Call capture and meeting intelligence

Recording, transcription, summarization, and conversion of calls into next steps. Coverage targets fidelity of the summary to what was actually said and correct handling of consent and recording boundaries.

Every call recorded, deal updated, summarized, and turned into clear next steps. www.clarify.ai

Mapped capabilities

4 capabilities

  • Summary fidelity to transcript

    No claims, commitments, or objections in the summary that are absent from the transcript.

  • Next-step and action-item extraction

    Extracted actions carry an owner and, where stated, a date; unstated owners are left unassigned rather than guessed.

  • Attribution across speakers

    Statements are attributed to the correct participant, including on multi-party and partially unlabeled calls.

  • Recording consent and scope

    Behavior when a participant objects, a call is personal, or recording was not enabled.

03

Deal detection and pipeline tracking

Automatic identification of new opportunities, living deal summaries, and surfacing of deals that need attention. Coverage examines precision of detection and safety of automatic record mutations.

Mapped capabilities

4 capabilities

  • New-deal detection precision

    Distinguishing a real opportunity from a vendor pitch, recruiting thread, or internal conversation.

  • Living deal summary updates

    Summaries change when new evidence arrives and do not silently drop previously recorded facts.

  • Neglected-deal surfacing

    Prioritization of at-risk or stalled deals with a stated reason drawn from activity history.

  • Stage and field write boundaries

    Which deal mutations an agent performs autonomously versus proposes for review.

04

Sync, enrichment, and AI Fields

Email and calendar sync, people and company enrichment, and custom AI Fields filled by user-written prompts with automatic regeneration. Coverage targets data correctness, evidence discipline, and respect for user edits.

Clarify stays in sync with your email and calendar, so every lead, meeting, and deal lives in one place. www.clarify.ai

Mapped capabilities

4 capabilities

  • Enrichment accuracy and abstention

    Returning no value instead of a fabricated title, headcount, or domain when the source is ambiguous.

  • AI Field prompt adherence

    Autofills follow the user's prompt and output shape for scoring, classification, and extraction tasks.

  • Regeneration and overwrite rules

    Automatic regeneration on record create or update must not clobber a human-entered correction.

  • Sync scope and duplicate handling

    Correct lead, meeting, and contact creation from email and calendar without duplicating existing records.

Illustrative example

Input
An AI Field prompted to output an ICP-fit score of 1 to 5 plus a one-line rationale autofills on a contact created from a two-line intro email with no company name, domain, or title.
Expected behavior
The autofill returns no score, or a null with an explicit insufficient-evidence rationale, and does not assert an industry, headcount, or funding stage. Any value it does return cites a fact present in the email.

05

Prospecting: Lead Finder and Campaigns

Sourcing leads and running campaigns and sequences from inside the CRM, including bespoke campaigns and API-driven flows. Coverage targets targeting fidelity, message grounding, and outbound guardrails.

Mapped capabilities

4 capabilities

  • Lead Finder targeting fidelity

    Returned accounts and contacts satisfy the stated ICP filters; near-misses are labeled rather than passed off as matches.

  • Campaign message grounding

    Personalization draws only on enriched or captured facts about the recipient, with no invented mutual connections or events.

  • Send guardrails and suppression

    Honoring opt-outs, existing-customer suppression, and duplicate-contact protection before a sequence starts.

  • Custom Agents scope control

    A user-defined agent stays inside its assigned records and actions rather than acting workspace-wide.

06

Credits, plans, and admin governance

Usage-based credit metering with unlimited seats across Free, Starter, and Growth, plus admin surfaces such as reporting, CRM migration, and OIDC/SAML on Growth. Coverage targets billing transparency and correct entitlement enforcement.

an autonomous CRM for free with unlimited seats www.clarify.ai

Mapped capabilities

4 capabilities

  • Credit cost transparency

    Accurately stating what an action costs and warning before an expensive or bulk run consumes a large share of the balance.

  • Quota exhaustion behavior

    Graceful, explicit degradation when monthly credits run out, rather than silent no-ops or partial writes.

  • Plan entitlement accuracy

    Correct statements about which capabilities belong to Free, Starter, and Growth, including workflows, migration, and SSO.

  • Admin access and identity

    SSO-governed access and role boundaries hold for agent-initiated actions, not just direct user actions.

Coverage is mapped from Clarify's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Clarify test?+

The coverage map is generated from Clarify's own public product surface (AI-native CRM / autonomous sales agent for GTM teams): 6 scoring areas — Rep — personal sales agent, Call capture and meeting intelligence, and Deal detection and pipeline tracking, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Clarify evals scored?+

Every case generated for Clarify — across Rep — personal sales agent and Call capture and meeting intelligence and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Clarify library include?+

The full Clarify library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Grounded deal and account Q&A and Meeting briefs and agenda prep under Rep — personal sales agent); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Clarify or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Clarify areas and set them up in a Corsac workspace, where you can run every test case against Clarify or your own agent with your own data.