All evals
F

Eval directory

Evals for Floworks

Eval coverage for Floworks, mapped from its public product surface.

About Floworks

Floworks AI sells a team of autonomous AI sales agents — Jesse (prospect research), Alisha (personalized cold email), Linda (LinkedIn outreach), and Sam (AI phone SDR) — that research prospects, run multi-channel outreach, handle objections, and book meetings. The agents are built on the company's ThorV2 model and pull from a 700M+ contact database plus web-wide intent signals. It is sold in monthly/annual tiers (Pro pay-per-meeting, Ultra at $498/month, and an Advanced enterprise plan) with CRM integration and email warmup.

Industry

AI SDR agents for B2B sales outreach

Use the eval library for Floworks

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Floworks?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Prospect Research & ICP Targeting (Jesse)

Turning a stated ideal customer profile into an enriched, verified prospect list drawn from the 700M+ contact database and web-wide sources, with buying-signal detection.

Jesse is an internet research agent that searches the web and a database of 700M+ contacts using intent data www.floworks.ai

Mapped capabilities

4 capabilities

  • ICP definition to search criteria

    Translating industry, size, geography, tech stack, and signal requirements into a coherent targeting query.

  • Web-wide and database discovery

    Sourcing beyond standard contact databases, including directories, marketplaces, and niche platforms.

  • Intent signal detection

    Identifying hiring activity, funding, technology adoption, pain-point mentions, and competitor-switching signals.

  • List enrichment and contact verification

    Delivering clean records with verified contacts and company context attached.

02

Cold Email Personalization & Reply Handling (Alisha)

Drafting research-grounded personalized email, sustaining the thread through objections, and converting replies into booked meetings.

Books qualified meetings directly on your calendar during the call without back-and-forth scheduling. www.floworks.ai

Mapped capabilities

4 capabilities

  • Research-grounded personalization

    Composing openers and value framing from prospect-specific data points rather than templates.

  • Brand voice and tone alignment

    Keeping generated copy consistent with the customer's stated positioning and tone.

  • Objection classification and response

    Recognizing objection type and replying in a way that keeps the conversation moving.

  • Multi-stakeholder scheduling

    Resolving calendar back-and-forth across several participants to land a meeting.

03

LinkedIn Outreach Automation (Linda)

Running connection requests, personalized messaging, and multi-touch nurture on LinkedIn while respecting platform constraints and account safety.

Mapped capabilities

4 capabilities

  • Connection targeting and note personalization

    Using profile content, recent activity, and mutual connections to justify the request.

  • Multi-touch follow-up sequencing

    Managing timing and cadence from acceptance through message sequence to meeting.

  • Account safety and send limits

    Staying within daily activity limits and platform norms to avoid account restriction.

  • Engagement-based prioritization

    Ranking prospects by likelihood to accept and engage.

04

Voice SDR Calling (Sam)

Outbound phone prospecting: script-driven qualification calls that adapt live, handle objections, and book on the calendar, with reviewable artifacts afterward.

Makes phone calls with human-like speech patterns, tone, and pacing that prospects can't distinguish from real people. www.floworks.ai

Mapped capabilities

4 capabilities

  • Script and qualification adherence

    Following the configured script and asking the customer's qualification questions.

  • Live objection handling

    Adapting to conversation flow and answering questions in real time.

  • In-call meeting booking

    Booking directly on the calendar during the call without follow-up scheduling.

  • Call artifacts and CRM sync

    Producing transcripts, recordings, and analytics, and routing qualified leads to CRM.

05

Factual Grounding & Claim Discipline (ThorV2)

The anti-hallucination surface: everything an agent asserts about a prospect, a company, or Floworks itself should trace to retrieved data, with graceful behavior when research is thin.

Mapped capabilities

4 capabilities

  • Prospect fact attribution

    Grounding every company or person claim in a retrieved source or enrichment field.

  • Thin-research handling

    Degrading to generic-but-honest outreach instead of inventing detail when signals are missing.

  • Stale or conflicting signal handling

    Behavior when enrichment records disagree or are visibly out of date.

  • Self-description accuracy

    Describing Floworks' own capabilities and metrics without overstatement.

Illustrative example

Input
Write Alisha's opening line for this prospect. Enrichment record: name, title, company, industry. No funding, hiring, news, or technology fields are populated.
Expected behavior
The opener uses only the four populated fields, or states that research is insufficient for a personalized hook. It does not assert a funding round, hiring push, product launch, or tooling change that the record does not contain.

06

Plans, Entitlements & Deliverability Operations

Correctly representing tier boundaries (Pro pay-per-meeting, Ultra $498/month, Advanced) and operating the sending infrastructure — warmup, inbox connection, CRM integration.

Mapped capabilities

4 capabilities

  • Tier and quota boundaries

    Contact counts, reachout volumes, and which features belong to which plan.

  • Billing model explanation

    Distinguishing pay-per-meeting from subscription and monthly from annual terms.

  • Email warmup and deliverability

    Inbox connection, warmup behavior, and sending-reputation protection.

  • CRM integration behavior

    Syncing prospects, replies, and booked meetings into the connected CRM.

Illustrative example

Input
I just signed up for Ultra at $498 billed monthly. Can I connect my CRM this month?
Expected behavior
Answers no for a monthly Ultra subscription, states that CRM integration is available on annual plans only, and points to switching to annual billing as the path to enable it.

Coverage is mapped from Floworks's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Floworks test?+

The coverage map is generated from Floworks's own public product surface (AI SDR agents for B2B sales outreach): 6 scoring areas — Prospect Research & ICP Targeting (Jesse), Cold Email Personalization & Reply Handling (Alisha), and LinkedIn Outreach Automation (Linda), and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Floworks evals scored?+

Every case generated for Floworks — across Prospect Research & ICP Targeting (Jesse) and Cold Email Personalization & Reply Handling (Alisha) and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Floworks library include?+

The full Floworks library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, ICP definition to search criteria and Web-wide and database discovery under Prospect Research & ICP Targeting (Jesse)); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Floworks or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Floworks areas and set them up in a Corsac workspace, where you can run every test case against Floworks or your own agent with your own data.