All evals
L

Eval directory

Evals for Leadbay

Eval coverage for Leadbay, mapped from its public product surface.

About Leadbay

Leadbay is an AI product that discovers and qualifies B2B leads, aimed at sales reps working fragmented, data-scarce markets. It delivers pre-qualified leads inside its own app and inside AI assistants such as Claude, ChatGPT, and Copilot via an MCP integration that includes its own UX/UI component library. The system is trained to imitate top reps' intuition and learns from market data, internal systems (CRM, ERP, CSVs), and user feedback.

Industry

B2B lead discovery and qualification AI (agentic prospecting)

Website

leadbay.ai

Use the eval library for Leadbay

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Leadbay?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Lead Discovery & Qualification

The core capability: searching, reasoning, and extrapolating over data-scarce markets to surface and score accounts the way a top rep would, including accounts with almost no online presence.

Leadbay searches, reasons, extrapolates, and qualifies the most data-scarce markets - the way your best reps would. leadbay.ai

Mapped capabilities

4 capabilities

  • Discovery in low-signal markets

    Surfacing candidate accounts that have minimal or no online footprint, rather than defaulting to whatever is easily indexed.

  • Qualification reasoning

    Producing a defensible rationale per account — why this lead fits — beyond filter matching or signal scraping.

  • Dynamic scoring and freshness

    Scoring leads on the described dynamic/enriched basis and reflecting when underlying data is stale.

  • Hard-to-find attribute questions

    Handling the attribute questions called out in the manifesto: seasonal hiring, CRM presence, expansion — including when they are not answerable.

Illustrative example

Input
In my Houston territory, find mechanical and HVAC contractors that look like they're expanding into commercial work. Most of them barely have a website.
Expected behavior
Returns accounts scoped to the Houston territory, each with a stated qualification rationale. Attributes that could not be sourced — such as expansion — are presented as inferences with their basis, not as confirmed facts.

02

Cold-Start Personalization from Customer Data

Leadbay claims leads are already qualified at first login, built from the customer's website and past won/lost deals, plus internal systems (CRM, ERP, CSVs).

Your leads are already qualified, built from your website and your past won/lost deals. leadbay.ai

Mapped capabilities

4 capabilities

  • ICP inference from website and deal history

    Deriving a usable target profile from a customer's site and won/lost outcomes without an explicit filter setup.

  • Internal system ingestion

    Reading CRM, ERP, and CSV inputs, including messy, partial, or conflicting records.

  • First-session usefulness

    Whether the zero-configuration first view is genuinely populated and relevant, not empty or generic.

  • Territory and market-cluster scoping

    Respecting a rep's assigned territory or named cluster when proposing accounts.

03

Feedback Learning Loop

The stated mechanism by which like / dislike / contact actions and frustrations expressed in chat push subsequent recommendations toward better prospects.

Leadbay learns what your best wins look like, then uses every frustration you share in chat leadbay.ai

Mapped capabilities

4 capabilities

  • Explicit signal uptake

    Like, dislike, and contact actions measurably shifting later lead sets in the indicated direction.

  • Conversational correction

    Turning a complaint stated in chat into a durable change rather than a one-turn fix.

  • Stability under sparse feedback

    Avoiding overreaction to a single rejection or an unrepresentative early signal.

  • Preference conflict handling

    Behavior when new feedback contradicts the profile inferred from won/lost history.

04

MCP Delivery & UX/UI Component Library

Leadbay ships as an MCP integration into Claude, ChatGPT, and Copilot, including its own library of UX/UI components used to build context-shaped apps and widgets.

The Leadbay MCP ships a full library of UX/UI components leadbay.ai

Mapped capabilities

4 capabilities

  • Tool and capability exposure

    What the MCP surface advertises to a host assistant and whether calls behave as described.

  • Component selection and rendering

    Choosing an appropriate widget from the component library for a given request and returning it well-formed.

  • Cross-host consistency

    Comparable behavior across the three named hosts — Claude, ChatGPT, and Copilot.

  • Conversational refinement in-host

    Refining or re-scoping a lead set from natural-language follow-ups inside the assistant.

05

Manager Visibility (Pre-CRM Activity)

The Manager Dashboard positions itself as a pre-CRM layer that makes exploration and prospecting activity — not just booked meetings — measurable and actionable.

Mapped capabilities

4 capabilities

  • Activity vs. outcome reporting

    Distinguishing exploration and research effort from end results such as qualified meetings.

  • Actionable KPI framing

    Reporting the simple, actionable indicators managers asked for instead of raw result counts.

  • Coverage of the pre-CRM gap

    Representing accounts a rep explored that do not yet exist as CRM records.

  • Team and territory rollup

    Aggregating across reps and territories without misattributing activity.

06

Grounding & Claim Discipline

Because the product asserts facts about companies with little online presence, and operates bilingually (EN/FR) across European and U.S. markets, separating evidence from inference is load-bearing.

It spots the invisible. Even with almost no online presence. leadbay.ai

Mapped capabilities

4 capabilities

  • Evidence vs. inference labeling

    Marking extrapolated attributes as inferred rather than presenting them as verified fact.

  • Refusal to fabricate firmographics

    Declining to invent contacts, vendors, headcount, or financials when no source supports them.

  • Provenance on qualification claims

    Attaching a traceable basis to the claims that drive a lead's score.

  • Bilingual and cross-market handling

    Consistent behavior on French and English inputs and on EU vs. U.S. company data.

Illustrative example

Input
Does AiRCO Mechanical use a CRM, and do they hire seasonally?
Expected behavior
States plainly that neither attribute is confirmed from available sources. Any leaning is offered as an inference with its reasoning, and no specific CRM vendor or hiring pattern is asserted as established fact.

Coverage is mapped from Leadbay's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Leadbay test?+

The coverage map is generated from Leadbay's own public product surface (B2B lead discovery and qualification AI (agentic prospecting)): 6 scoring areas — Lead Discovery & Qualification, Cold-Start Personalization from Customer Data, and Feedback Learning Loop, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Leadbay evals scored?+

Every case generated for Leadbay — across Lead Discovery & Qualification and Cold-Start Personalization from Customer Data and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Leadbay library include?+

The full Leadbay library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Discovery in low-signal markets and Qualification reasoning under Lead Discovery & Qualification); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Leadbay or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Leadbay areas and set them up in a Corsac workspace, where you can run every test case against Leadbay or your own agent with your own data.