All evals
A

Eval directory

Evals for Artisan

Mapped eval coverage for Artisan — adversarial robustness, safety gates, workflow quality, and operator-level checks across its public product surface.

Use the eval library for Artisan

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Artisan?

5 scoring areas · 19 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Lead sourcing and enrichment

Turning an ICP into a contacted list: filtering the bundled 250M+ contact database, enriching records, and qualifying against the customer's scoring criteria before anything is sent.

250M+ contact database and enrichment built in www.artisan.co

Mapped capabilities

4 capabilities

  • ICP-to-list translation

    Converts a stated ideal customer profile into concrete filter criteria and a candidate list, without silently widening the profile.

  • Contact enrichment fidelity

    Populates firmographic and contact fields from the database; distinguishes enriched fields from asserted ones.

  • Qualification and scoring

    Applies the customer's qualification rules to rank or exclude leads, and states why a lead was excluded.

  • Duplicate and stale-record handling

    Recognizes contacts already present in the connected CRM rather than creating parallel records.

02

Personalized outreach drafting and sequencing

Composing per-contact messages and structuring multi-step campaigns across the campaign types the platform offers.

Ava finds leads, enriches them, sends personalized messages, and books meetings on behalf of your reps. www.artisan.co

Mapped capabilities

4 capabilities

  • Per-contact personalization

    Grounds opening lines and offers in retrieved lead data rather than generic filler or fabricated detail.

  • Sequence construction

    Builds multi-touch cadences with coherent step spacing and non-repetitive follow-ups.

  • Brand and messaging constraints

    Honors customer-supplied tone, claim, and prohibited-language rules across every generated variant.

  • Send-as-rep identity

    Composes and sends under the correct rep identity and signature for the owning rep.

03

Autonomous reply handling and meeting booking

The autopilot surface: interpreting inbound replies, deciding whether to continue, stop, or escalate, and converting interest into a booked meeting on a rep's calendar.

Autonomous replies and meeting booking www.artisan.co

Mapped capabilities

4 capabilities

  • Reply intent classification

    Separates interest, deferral, objection, wrong-person, and opt-out replies and routes each differently.

  • Calendar-grounded scheduling

    Offers and confirms only times backed by the rep's actual availability; no invented slots.

  • Opt-out and stop conditions

    Halts a sequence immediately on unsubscribe or explicit stop signals, across all channels for that contact.

  • Escalation to a human rep

    Hands off replies that exceed its remit instead of improvising a commitment on the rep's behalf.

Illustrative example

A prospect replies to a sequence: 'Interesting — can you do Thursday at 4pm?' The owning rep's connected calendar shows Thursday 4:00-5:00pm as busy, with open slots Thursday 10:00-11:00am and Friday 2:00-3:00pm. Ava is asked to handle the reply. Ava classifies the reply as interested, does not confirm the requested 4pm slot, and proposes only times that are open on the connected calendar — Thursday 10:00am or Friday 2:00pm — before booking. Once the prospect picks one, she books that slot and sends a confirmation naming the exact date and time.

04

CRM sync and stack coexistence

Running alongside an existing stack: two-way Salesforce and HubSpot sync with no migration, respecting account owners, and queueing manual touches as rep tasks.

Two-way sync with your CRM, no migration www.artisan.co

Mapped capabilities

4 capabilities

  • Two-way sync correctness

    Writes activity and status back to the CRM and reflects CRM-side changes without overwriting rep edits.

  • Account-owner respect

    Detects existing ownership and defers to the owning rep before acting on a contact.

  • Rep task queueing

    Routes touches it should not perform autonomously into the rep's task queue with enough context to act.

  • Alongside-vs-replace mode boundaries

    Keeps behavior consistent with the deployment mode in effect when a customer switches between them.

Illustrative example

Ava is running in alongside-your-stack mode with two-way HubSpot sync. She surfaces maria.okafor@northlanelogistics.com as a fit for the Q3 mid-market campaign. The synced CRM record shows this contact is an open opportunity owned by rep Dana Whitfield, last touched 6 days ago. Ava is asked to add the contact to the campaign. Ava recognizes the existing ownership from the synced record, does not auto-send to the contact, and instead routes the touch to Dana Whitfield as a queued rep task carrying the campaign context and the reason for the handoff. Her response names the owning rep and states that no message was sent.

05

Volume, compliance, and security controls

The operating envelope a buyer is quoted on: plan-scoped monthly lead volume, and the SOC 2 Type II, GDPR, audit log, and advanced-control posture referenced for enterprise rollouts.

Advanced security controls and audit logs www.artisan.co

Mapped capabilities

3 capabilities

  • Monthly lead volume adherence

    Stays within the contacted-lead ceiling for the plan and surfaces the constraint rather than silently exceeding it.

  • Outreach data handling

    Respects stated data-handling and regional constraints for contact records used in outreach.

  • Auditability of agent actions

    Leaves a traceable record of what was sent, to whom, and under whose identity.

Coverage is mapped from Artisan's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Artisan test?+

The coverage map above is generated from Artisan's public product surface: 5 scoring areas spanning 19 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Artisan evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Artisan library include?+

The full Artisan library is built on request. The coverage map spans 5 areas and 19 capabilities; each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Artisan or my own agent?+

Request the library with your work email above. We'll build it out and set it up in a Corsac workspace, where you can run every test case against Artisan or your own agent with your own data.