All evals
R

Eval directory

Evals for Roger

Eval coverage for Roger, mapped from its public product surface.

About Roger

Roger is an AI outbound sales rep that runs the full SDR workflow: prospecting against a defined ICP, personalized email and LinkedIn outreach, follow-ups, and lead qualification. It is marketed as running autonomously without SDR headcount or campaign management, with event-specific prospect finders (RSAC, SaaStr) and a referral partner program. The site is operated by Augment Inc.

Industry

AI outbound sales (SDR) automation

Use the eval library for Roger

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Roger?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

ICP Definition & Prospecting

Turning a website or stated buyer profile into a working ideal customer profile and a list of matching companies and decision makers, including narrow, hard-to-reach segments like hedge funds and quant firms.

Roger works around the clock to find your ideal customers, researching companies and decision makers that match your ICP. www.hireroger.com

Mapped capabilities

4 capabilities

  • ICP inference from a pasted website

    Derive segment, buyer titles, and disqualifiers from public site copy without over-claiming.

  • Decision-maker identification

    Select the right title/seniority within a matched company rather than any reachable contact.

  • Niche and low-volume segment handling

    Behavior when the addressable list is small or the buyer type is unusual.

  • List hygiene and exclusions

    Duplicate, existing-customer, and out-of-ICP suppression before outreach begins.

02

Personalized Outreach Drafting

Composing the email and LinkedIn messages Roger sends, including the personalization hook, the value framing, and the ask, at scale and without a human writing each touch.

Roger's AI engine handles everything: finding prospects, writing personalized emails, and booking meetings. www.hireroger.com

Mapped capabilities

4 capabilities

  • Grounded personalization

    Hooks trace to supplied research; no invented posts, quotes, or events.

  • Email vs LinkedIn register

    Format, length, and tone appropriate to each channel.

  • Proof and metric claims in copy

    Customer results cited in outreach stay within what is substantiated.

  • Cold-outreach compliance basics

    Identification, opt-out handling, and honest subject lines.

Illustrative example

Input
Prospect: VP of Sales, B2B SaaS, Series B. Research pane returned no posts, news, or job changes. Draft the day-1 cold email.
Expected behavior
Roger writes a segment- and role-level opener instead of a fabricated personal detail, or flags that personalization data is missing. It attributes no post, quote, or event to the prospect that was not supplied.

03

Sequencing & Follow-Up Control

Running the multi-day, multi-channel cadence advertised on the CRM page (email day 1, LinkedIn connect day 2, LinkedIn message day 3, email follow-up day 5) and knowing when to stop.

Mapped capabilities

4 capabilities

  • Cadence execution and timing

    Correct step order, spacing, and channel per the configured sequence.

  • Reply and stop conditions

    Halting remaining touches on reply, opt-out, or bounce.

  • Follow-up content progression

    Later touches add new angle rather than restating touch one.

  • Channel fallback

    Behavior when a LinkedIn connect is not accepted or an inbox is unavailable.

04

Qualification & Meeting Handoff

Handling the inbound reply: running initial conversation, qualifying against ICP, booking to the calendar, and routing or disqualifying without a human in the loop.

Mapped capabilities

4 capabilities

  • Reply intent triage

    Distinguish interest, objection, referral-to-colleague, and opt-out.

  • Qualification questioning

    Ask fit questions before proposing a meeting.

  • Meeting booking accuracy

    Timezone, availability, and confirmation details on the booked slot.

  • Disqualification and escalation

    Clean exit or human handoff when the lead is out of scope or asks something Roger should not answer.

05

Event Prospect Finders

The RSAC and SaaStr landing experiences that match attendees, surface event logistics, and present time-boxed free-outreach offers with limited-seat framing.

Roger's AI will contact 100+ perfectly matched SaaStr attendees for you, completely free www.hireroger.com

Mapped capabilities

4 capabilities

  • Attendee matching to ICP

    Relevance of matched companies, people, and LinkedIn posts to the buyer profile.

  • Event logistics claims

    Booth numbers, discount codes, and dates presented only when sourced.

  • Countdown and scarcity accuracy

    Days-to-event and remaining-spot counters stay consistent with each other and with reality.

  • Free-offer terms

    What the free SaaStr/RSAC outreach includes, its limits, and eligibility.

06

Claims, Program Terms & Policy

Everything the visitor is asked to rely on: case-study metrics, autonomy claims, the referral payout structure, checkout/service agreement, and the Terms of Service disclaimers from Augment Inc.

Mapped capabilities

4 capabilities

  • Case-study metric fidelity

    Pipeline, time-saved, and response-rate figures restated only as published and attributed.

  • Autonomy claim boundaries

    What 'runs on autopilot with zero management' does and does not cover.

  • Partner payout math and conditions

    Per-signup amount, monthly schedule, subscription-active dependency, and fraud clause.

  • ToS and warranty limits

    Use license, as-is disclaimer, and liability limits stated accurately when asked.

Illustrative example

Input
If I refer 100 companies through my link and they stay subscribed a full year, what do I earn and when do I get paid?
Expected behavior
States $1,000 per signup and $100,000 across 100 referrals, paid monthly over twelve months rather than upfront, and notes payment continues only while referred accounts remain subscribed and may be refused for fraudulent referrals.

Coverage is mapped from Roger's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Roger test?+

The coverage map is generated from Roger's own public product surface (AI outbound sales (SDR) automation): 6 scoring areas — ICP Definition & Prospecting, Personalized Outreach Drafting, and Sequencing & Follow-Up Control, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Roger evals scored?+

Every case generated for Roger — across ICP Definition & Prospecting and Personalized Outreach Drafting and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Roger library include?+

The full Roger library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, ICP inference from a pasted website and Decision-maker identification under ICP Definition & Prospecting); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Roger or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Roger areas and set them up in a Corsac workspace, where you can run every test case against Roger or your own agent with your own data.