All evals
Artisan

Eval directory

Evals for Artisan

Eval coverage for Artisan, mapped from its public product surface.

About Artisan

Artisan sells Ava, an AI business development representative that runs outbound sales motions on behalf of a team's reps. Ava finds and enriches leads, sends personalized messages, handles replies, and books meetings, either plugged into an existing sales stack or replacing it as an end-to-end outbound platform. Plans are quoted by contacted-lead volume (Team, Scale, Enterprise) and include CRM sync, a 250M+ contact database, CSM onboarding, and an AI dialer add-on.

Industry

AI sales development / outbound BDR automation

Use the eval library for Artisan

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Artisan?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Lead Sourcing & Enrichment

Turning an ideal customer profile into a qualified, enriched contact list drawn from the 250M+ B2B contact database.

“250M+ contact database and enrichment built in” www.artisan.co

Mapped capabilities

4 capabilities

  • ICP-based lead filtering

    Translating a stated ideal customer profile into concrete filters and a candidate list.

  • Lead qualification and scoring

    Applying the team's scoring criteria to rank or exclude candidates.

  • Contact enrichment coverage

    Populating firmographic and contact fields, and flagging fields that could not be verified.

  • Database claim discipline

    Describing coverage and verification of the contact database without overstating it.

02

Personalized Outbound Messaging

Drafting and sequencing personalized outbound touches on behalf of a named rep, across the supported campaign types.

“Ava finds leads, enriches them, sends personalized messages, and books meetings on behalf of your reps.” www.artisan.co

Mapped capabilities

4 capabilities

  • Personalization grounded in enriched data

    Using only researched or enriched facts about the prospect in the message body.

  • Campaign type and sequence selection

    Choosing an appropriate campaign shape and follow-up cadence for the stated motion.

  • Sends-as-rep voice and signature

    Matching the sending rep's identity, tone, and signature conventions.

  • Prospect fact discipline

    Declining to assert unverified details about a prospect or their company.

03

Reply Handling & Meeting Booking

Reading inbound replies, deciding what to do next, and either booking a meeting or handing off to a human.

“Autonomous replies and meeting booking” www.artisan.co

Mapped capabilities

4 capabilities

  • Reply intent classification

    Separating interest, objection, referral, out-of-office, and opt-out.

  • Autonomous reply drafting

    Responding in-thread to common objections and questions.

  • Meeting scheduling and calendar handoff

    Proposing times and confirming a booking against the right rep's calendar.

  • Escalation to a human rep

    Stopping automation and queuing the thread when the reply exceeds Ava's remit.

Illustrative example

Input
Prospect replies to a sequence: "Wrong person for this — you want Priya, our VP of Operations. Copying her here."
Expected behavior
Ava should treat this as a referral, not interest: stop the sequence for the original contact, capture Priya as the new lead for the account, and hand the thread to the rep rather than proposing meeting times to the person who declined.

04

CRM Sync & Stack Integration

Running inside an existing sales stack: two-way CRM sync, ownership rules, and the boundary between Ava's work and the rep's.

Mapped capabilities

4 capabilities

  • Two-way Salesforce and HubSpot sync

    Writing activity back and reading records without requiring a migration.

  • Account ownership rules

    Respecting existing account owners before contacting or routing a lead.

  • Manual touches as rep tasks

    Queuing work Ava should not do autonomously as a task for the rep.

  • Alongside-stack vs. full-platform mode

    Behaving consistently with whether Ava is plugged in or running the whole motion.

05

Plans, Volume & Commercial Scoping

Answering pricing and packaging questions across Team, Scale, and Enterprise, which are quoted by contacted-lead volume.

“Trusted by 6,000+ sales teams, from startups to enterprise” www.artisan.co

Mapped capabilities

4 capabilities

  • Volume-to-tier mapping

    Matching a stated monthly contacted-lead volume to the right plan tier.

  • Plan inclusion questions

    Stating what every plan includes versus what is Scale- or Enterprise-only.

  • AI dialer add-on

    Describing the dialer as a per-seat add-on rather than bundled capacity.

  • Routing to sales for quotes

    Deferring to a scoped conversation instead of inventing a price.

Illustrative example

Input
"We're contacting around 6,000 leads a month and want two reps on the dialer. What will this cost us monthly?"
Expected behavior
Ava should map 6,000 contacted leads per month to the Scale plan, note that the AI dialer is a per-seat add-on for the two reps, and state that pricing is scoped with sales rather than quoting a monthly figure.

06

Trust, Compliance & Rollout Support

Statements about security posture, data handling, and the human support that ships with each plan.

“Advanced security controls and audit logs” www.artisan.co

Mapped capabilities

4 capabilities

  • SOC 2 Type II and GDPR posture

    Representing certifications and data-handling commitments accurately.

  • Enterprise security controls and audit logs

    Scoping advanced controls to the plans that actually carry them.

  • CSM onboarding scope

    Describing what guided onboarding and shared-channel support cover.

  • Forward-deployed strategist scope

    Limiting strategist buildout claims to Enterprise rollouts.

Coverage is mapped from Artisan's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Artisan test?+

The coverage map is generated from Artisan's own public product surface (AI sales development / outbound BDR automation): 6 scoring areas — Lead Sourcing & Enrichment, Personalized Outbound Messaging, and Reply Handling & Meeting Booking, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Artisan evals scored?+

Every case generated for Artisan — across Lead Sourcing & Enrichment and Personalized Outbound Messaging and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Artisan library include?+

The full Artisan library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, ICP-based lead filtering and Lead qualification and scoring under Lead Sourcing & Enrichment); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Artisan or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Artisan areas and set them up in a Corsac workspace, where you can run every test case against Artisan or your own agent with your own data.