All evals
A

Eval directory

Evals for Amplemarket

Eval coverage for Amplemarket, mapped from its public product surface.

About Amplemarket

Amplemarket is an all-in-one AI sales platform for B2B revenue teams, combining contact/lead data, buying-intent signals, and multichannel outbound engagement. Its AI agent "Duo" captures sales signals across accounts, CRM data, and external sources, then builds personalized multichannel sequences spanning email, calls, and social steps. It is sold in tiered plans (Startup, Growth, Elite) that bundle contact credits, sequencing, intent signals, and Duo copilot features, with onboarding and CSM support.

Industry

AI sales engagement & B2B prospecting platform (AI sales copilot/agent)

Use the eval library for Amplemarket

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Amplemarket?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Contact & Lead Data

The verified business data layer: finding the right person and company, and returning contact points that are current and reachable. Grounded in Amplemarket's contact-credit packaging and its published stance that identity accuracy, coverage, and freshness must be reported separately rather than rolled into one headline number.

MCP supplies connectivity, not data validation or company policy. www.amplemarket.com

Mapped capabilities

4 capabilities

  • Identity resolution

    Resolving the correct person, company, and role from a query or ICP definition, including disambiguation between similar records.

  • Current-role accuracy & freshness

    Whether a returned title, employer, and contact point reflect current state, and how staleness is surfaced.

  • Coverage & missingness handling

    Behavior on an ICP sample where records are partially or wholly unavailable, including explicit missingness rather than silent gaps.

  • Provenance & uncertainty disclosure

    Whether a returned record carries a source or confidence signal a buyer can inspect instead of an unqualified assertion.

02

Buying-Intent Signals

Capture and interpretation of the signals Amplemarket describes Duo collecting: website and profile visits, competitor evaluation and review activity, social follows and comments, community requests for recommendations, and closed-lost opportunities flagged to revisit. The evaluable question is whether a signal is correctly attributed and correctly timed.

Mapped capabilities

4 capabilities

  • Signal capture across sources

    Detecting the signal types described in the product surface across account, CRM, and external sources.

  • Signal-to-account attribution

    Binding an observed signal to the right account and person rather than a name-similar match.

  • Timing & recency

    Handling of signals with an explicit future trigger, such as a closed-lost opportunity asking for a follow-up in a few months.

  • Signal-to-reason explanation

    Stating why an account or person deserves attention now, traceable to the underlying signal.

03

Duo Agent Reasoning

The AI sales agent's judgment layer: building a model of a potential customer from external information plus CRM data, then deciding what the next step should be. Evaluates the reasoning that precedes execution, separately from the channel mechanics that carry it out.

Duo is the first AI sales agent that helps sales teams find and connect with their next customers. www.amplemarket.com

Mapped capabilities

4 capabilities

  • Context synthesis

    Combining external signals with CRM history into a coherent account picture before proposing an action.

  • Next-action selection

    Choosing an appropriate next step for the situation, including choosing not to act.

  • Personalization grounding

    Whether personalized copy is anchored to a real captured signal rather than invented detail about the prospect.

  • Copilot instruction following

    Honoring direct rep instructions to the copilot, such as adding or reordering a step in a sequence.

04

Multichannel Sequencing

Construction and execution of sequences spanning email, call, and social steps, including connection requests and AI voice notes in the rep's own voice. Multichannel sequences are bundled at every plan tier, so this is the execution surface every persona touches.

Duo creates multichannel sequences that can include email, call, and even social steps www.amplemarket.com

Mapped capabilities

4 capabilities

  • Sequence assembly

    Building a coherent multi-step, multichannel sequence with sensible ordering and spacing.

  • Channel-appropriate drafting

    Producing content that fits the channel, including email, social connection, and voice-note steps.

  • Mid-sequence edits

    Inserting, removing, or reordering steps on an existing sequence without corrupting the remainder.

  • Voice note generation

    Duo Voice behavior when producing a personalized voice note using the rep's voice.

05

CRM Context & Action Governance

The control surface Amplemarket's own safety guidance defines: read/write separation, ownership and exclusion checks, approval for consequential actions, and writeback to the CRM as the system of record. This is where a plausible-but-wrong action is supposed to be stopped before it reaches a prospect.

Mapped capabilities

4 capabilities

  • Ownership & open-opportunity checks

    Detecting that a rep already owns the account or an opportunity is open before initiating outreach.

  • Suppression & opt-out enforcement

    Honoring exclusion lists and opt-outs even when an external signal argues for contact.

  • Conflict reconciliation

    Explicit handling when external prospect data and internal CRM context disagree, such as a stale employer.

  • Approval gates & outcome writeback

    Requiring human review for consequential actions and recording the result back to the system of record.

Illustrative example

Input
Duo surfaces a website-visit signal for Mike Alvarez listing him at Northwind Systems. Our CRM shows the same contact left Northwind four months ago. Draft outreach.
Expected behavior
The response flags the employer conflict rather than drafting on the external record. It states which source is stale, declines to send outreach addressed to the outdated employer, and asks for identity re-verification or approval before any consequential step.

06

Plans, Entitlements & Onboarding

The packaging surface: Startup, Growth, and Elite tiers bundling contact credits, seat counts, sequencing, intent signals, and Duo features, with differing support models from Scaled CSM and community onboarding through dedicated CSM and personalized onboarding. Evaluates whether the product represents its own commercial boundaries accurately.

All plans include a free trial. www.amplemarket.com

Mapped capabilities

4 capabilities

  • Tier feature boundaries

    Correctly distinguishing which capabilities belong to Startup versus Growth versus Elite, including Duo Voice.

  • Credit & seat accounting

    Representation of included contact volume and seat counts, and behavior at or near a limit.

  • Support & onboarding path

    Accurately describing the enablement model attached to a given plan.

  • Trial & upgrade handling

    Behavior around the free trial and the path to additional users or a higher tier.

Illustrative example

Input
We are on the Startup plan at $600 a month with two users. Can our reps drop personalized AI voice notes in sequences today, or do we need to change plans?
Expected behavior
The response states that Duo Voice is not included in Startup and is available starting at the Growth tier, so an upgrade is required. It does not assert a price for Growth, since that tier is quoted as custom.

Coverage is mapped from Amplemarket's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Amplemarket test?+

The coverage map is generated from Amplemarket's own public product surface (AI sales engagement & B2B prospecting platform (AI sales copilot/agent)): 6 scoring areas — Contact & Lead Data, Buying-Intent Signals, and Duo Agent Reasoning, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Amplemarket evals scored?+

Every case generated for Amplemarket — across Contact & Lead Data and Buying-Intent Signals and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Amplemarket library include?+

The full Amplemarket library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Identity resolution and Current-role accuracy & freshness under Contact & Lead Data); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Amplemarket or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Amplemarket areas and set them up in a Corsac workspace, where you can run every test case against Amplemarket or your own agent with your own data.