All evals
O

Eval directory

Evals for OpenFunnel

Eval coverage for OpenFunnel, mapped from its public product surface.

About OpenFunnel

OpenFunnel provides headless, API-accessible primitives that coding agents and in-house GTM engineering teams can call to source and qualify target accounts. The documented primitives cover lookalike search from closed-won/closed-lost accounts, TAM building from an ICP definition, tech-stack search, deep account research, and hiring-initiative signals. The site positions the product for agent-driven go-to-market workflows and lists Y Combinator backing; several primitives are marked "coming soon" or gated behind "book a demo".

Industry

agent-facing GTM data & signal API (go-to-market primitives)

Use the eval library for OpenFunnel

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for OpenFunnel?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

02

TAM building from an ICP definition

The single-call TAM builder that takes an ICP definition and returns the addressable set; documented as one primitive currently gated behind "book a demo".

Build your full total addressable market from an ICP definition in one call. openfunnel.dev

Mapped capabilities

4 capabilities

  • ICP definition parsing

    Stated firmographic and qualitative ICP constraints are reflected in the returned set.

  • Constraint adherence

    Returned companies do not violate explicit exclusions in the ICP definition.

  • Single-call completeness

    One call yields a usable market set rather than requiring undocumented pagination assumptions.

  • Gated-access behavior

    Demo-gated status is surfaced clearly to a caller instead of failing opaquely.

04

Deep account research

Qualifying a single named account against a caller-supplied custom question via the documented POST deep-research primitive.

Deep-research and qualify a single account against a custom question. openfunnel.dev

Mapped capabilities

4 capabilities

  • Custom question qualification

    The returned judgment answers the caller's actual question about the account.

  • Evidence and citation

    Conclusions are accompanied by inspectable supporting sources.

  • Abstention on thin evidence

    Insufficient evidence yields an explicit non-answer rather than a confident guess.

  • Ambiguous account resolution

    Similarly named or renamed companies are disambiguated before research proceeds.

Illustrative example

Input
POST deep-research for a small private company with the custom question: does this account run an in-house GTM engineering team?
Expected behavior
Either a verdict accompanied by at least one inspectable source, or an explicit insufficient-evidence response. A confident yes or no with no cited source is a failure.

05

Hiring-initiative signals

Sourcing ICP companies from job-post signals — the five documented primitives covering role-based, named-initiative, pain-point/concept, and first-time-mention searches, all currently demo-gated.

Source ICP companies with active initiatives named in a single job post. openfunnel.dev

Mapped capabilities

4 capabilities

  • Role-based sourcing

    Companies hiring for a specified role are returned with the triggering post identifiable.

  • Named internal initiative

    An initiative named inside a job post is matched as a phrase, not a keyword smear.

  • Pain-point or concept mentions

    Conceptual matches in job posts are distinguished from literal keyword hits.

  • First-time-mention detection

    A first-ever mention of a role or concept is separated from ongoing recurring hiring.

06

Agent integration and data posture

How an agent onboards, discovers what is callable, and handles the documented mix of live endpoints, "coming soon" primitives, and demo-gated ones — plus the personal-data handling the published privacy policy governs.

Mapped capabilities

4 capabilities

  • Docs-to-first-call path

    An agent can go from the documented reference to a valid request without human clarification.

  • Availability semantics

    Coming-soon and demo-gated primitives are reported distinctly from real failures.

  • Bulk request behavior

    Bulk endpoints handle partial success without discarding usable results.

  • Personal-data handling

    People-level results stay consistent with the commitments in the published privacy policy.

Coverage is mapped from OpenFunnel's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for OpenFunnel test?+

The coverage map is generated from OpenFunnel's own public product surface (agent-facing GTM data & signal API (go-to-market primitives)): 6 scoring areas — Lookalike search, TAM building from an ICP definition, and Tech-stack search, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the OpenFunnel evals scored?+

Every case generated for OpenFunnel — across Lookalike search and TAM building from an ICP definition and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the OpenFunnel library include?+

The full OpenFunnel library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Closed-won seed expansion and Closed-lost handling under Lookalike search); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against OpenFunnel or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped OpenFunnel areas and set them up in a Corsac workspace, where you can run every test case against OpenFunnel or your own agent with your own data.