All evals
R

Eval directory

Evals for Regie.ai

Mapped eval coverage for Regie.ai — adversarial robustness, safety gates, workflow quality, and operator-level checks across its public product surface.

Use the eval library for Regie.ai

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Regie.ai?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agent–Rep Orchestration

The core RegieOne claim: task orchestration for AI agents and human reps in a single workflow, so plays execute consistently without reps guessing the next step.

Regie brings enrichment, dialing, email, sequencing, and reporting into one streamlined workflow so reps stay focused on selling. www.regie.ai

Mapped capabilities

4 capabilities

  • Task handoff between agent and rep

    Which steps an agent completes autonomously vs. routes to a rep, and whether the rep sees the full context on handoff.

  • Play execution completeness

    Every configured step in a play actually fires; no silently skipped or dropped touches.

  • Autonomous prospect discovery and enrollment

    Force Multiplier agents discovering and enrolling new prospects into an existing play without manual list upload.

  • Prompt-based custom steps

    Operator-authored prompt steps consolidating workflow actions and behaving as specified within a sequence.

Illustrative example

Enroll 50 contacts in a 6-step multi-channel play (email → call → social → email → call → email). During execution, force step 3 (social) to fail for 10 of the contacts with an upstream connector error. The 40 unaffected contacts complete all 6 steps in order. The 10 failed contacts surface the step-3 failure as an explicit, rep-visible error or task rather than being skipped silently, and either retry or hold at step 3 per the configured failure policy — no contact advances past a step recorded as failed.

02

Sequencing & Play Configuration

Building the exact multi-channel workflows a team wants — by persona, intent, or custom signal — across static and dynamic sequencing modes.

Mapped capabilities

4 capabilities

  • Static multi-channel sequencing

    Fixed-order email/call/social cadences executing on schedule across channels.

  • Intent-based dynamic sequencing

    Cadence branching or re-ordering in response to changing intent and engagement signals.

  • Persona and signal-based routing

    Correct play selection when a prospect matches a persona, an intent tier, or a custom signal definition.

  • Inbound and campaign-based agents

    Agents triggered by inbound activity or campaign membership rather than a manually built list.

03

Contact Data & Enrichment

Sourcing, enriching, and verifying account and contact records — the input quality layer behind every downstream touch, metered in AI + enrichment credits.

120,000 AI + enrichment credits (1k accounts / 4k contacts) included www.regie.ai

Mapped capabilities

4 capabilities

  • Custom enrichment waterfalls

    Provider fallback ordering, field precedence, and behavior when every source in the waterfall misses.

  • Bounce and job-change verification

    Flagging invalid addresses and stale titles before a contact is worked.

  • BYO list ingestion

    Customer-supplied lists mapped, deduplicated, and reconciled against enriched records.

  • Credit consumption and coverage limits

    Enrichment credit accounting against the stated account/contact allotments and behavior at exhaustion.

04

AI Messaging & Personalization

Generating persona-aligned, signal-grounded outreach copy and follow-ups — the claim that every touch is personalized rather than spray-and-pray.

Reach 3–5x more accounts with personalized, signal-driven outreach that reps simply don’t have time to do manually. www.regie.ai

Mapped capabilities

4 capabilities

  • Persona-aligned message generation

    Copy that reflects the targeted persona's role and priorities rather than generic templating.

  • Research agents and Reasons to Engage

    Personalization anchored to retrieved account research, with the cited reason traceable to a real source.

  • Merge field and variable integrity

    No unresolved placeholders, wrong-company mentions, or fabricated attributes in a sent message.

  • Follow-up thread coherence

    Sequenced follow-ups that reference prior touches without repeating or contradicting them.

Illustrative example

Generate a first-touch cold email to a VP of Revenue Operations at a mid-market logistics company, using only the enriched fields available on the record: company name, industry, headcount band, and one intent signal ("researching sales engagement software"). No research-agent findings are attached. The message personalizes on the supplied fields and the intent signal only. It does not assert a funding round, headline, executive quote, product launch, or named competitor that is absent from the provided record, and it does not fill gaps with plausible-sounding specifics.

05

Voice & Dialing

Built-in calling: the power dialer, parallel dialing across up to nine lines, Sales Floor, coaching, and voicemail drops.

Parallel dialer (up to 9 lines) www.regie.ai

Mapped capabilities

4 capabilities

  • Parallel dialer connect handling

    Routing a live answer to the rep while cleanly releasing the other simultaneous lines.

  • Power dialer queue and disposition

    Call outcomes writing back to the prospect record and advancing the sequence correctly.

  • AI voicemails and call coaching

    Voicemail drop accuracy and coaching feedback tied to the actual call transcript.

  • Sales Floor session behavior

    Shared dialing session state across participating reps.

06

Privacy, Data Handling & Integrations

Regie.ai's stated GDPR data-processor role — acting on documented controller instructions — plus the CRM/SEP sync and sending-infrastructure surfaces that carry customer personal data.

Regie.ai operates as a data processor, while its customers act as data controllers. www.regie.ai

Mapped capabilities

4 capabilities

  • Processor-scope adherence

    Processing bounded to the controller's documented instructions, without repurposing customer contact data.

  • Data subject rights assistance

    Supporting access, rectification, and erasure requests that reach the customer as controller.

  • CRM and SEP sync fidelity

    Bidirectional record sync with Salesforce, HubSpot, Dynamics, or Pipedrive without duplicate or overwritten fields.

  • Mailbox, domain, and warming controls

    Multi-mailbox sending, domain assignment, and warming schedules behaving as configured.

Coverage is mapped from Regie.ai's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Regie.ai test?+

The coverage map above is generated from Regie.ai's public product surface: 6 scoring areas spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Regie.ai evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Regie.ai library include?+

The full Regie.ai library is built on request. The coverage map spans 6 areas and 24 capabilities; each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Regie.ai or my own agent?+

Request the library with your work email above. We'll build it out and set it up in a Corsac workspace, where you can run every test case against Regie.ai or your own agent with your own data.