All evals
Scarf

Eval directory

Evals for Scarf

Eval coverage for Scarf, mapped from its public product surface.

About Scarf

Scarf is an open source usage intelligence platform that connects signals from package and container registries, documentation pixels, SDKs, and telemetry to show which companies are adopting a project and how their usage changes over time. It surfaces account-level adoption data that teams can query, filter, export, or sync into GTM tools via Slack, APIs, and webhooks. The product is positioned around privacy, with the company stating it does not retain personally identifiable information or raw IP addresses.

Industry

open source usage intelligence / GTM analytics

Use the eval library for Scarf

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Scarf?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Adoption Signal Capture

How Scarf ingests usage signals from the sources named in its public materials, and how a team connects them to a project.

Scarf does not charge for events, telemetry volume, package or download traffic, docs pixels, or CLI launch signals. about.scarf.sh

Mapped capabilities

4 capabilities

  • Package and container registry signals

    Downloads and pulls tracked via package registries and Scarf Gateway.

  • Documentation pixels

    Docs-page pixels as a discovery/evaluation signal.

  • SDK and telemetry signals

    SDK, CLI launch, and product telemetry events.

  • Custom event and data import

    Bringing your own event data; unlimited data imports.

02

Company-Level Adoption Intelligence

Turning raw signals into a view of which organizations are active, what they use, and how engagement changes over time.

Mapped capabilities

4 capabilities

  • Company identification and surfacing

    Resolving usage into named organizations and account records.

  • Version and component breakdown

    Which releases, packages, or components an account is using.

  • Adoption funnel stages

    Funnel stages spanning downloads, docs interaction, and ongoing usage.

  • Activity scoring customization

    Weighting signals so higher-intent activity ranks first.

03

Query, Filter, and Analysis

The self-serve surfaces for interrogating adoption data, including natural-language questions and programmatic access.

Connect package and container registries, Scarf Gateway, SDKs, pixels, or your own event data. about.scarf.sh

Mapped capabilities

4 capabilities

  • Natural-language questions in Scarf

    Asking direct questions of your usage data.

  • Real-time attribute filtering

    Filtering accounts by any attribute in the dataset.

  • Insights and historical charts

    Dashboards and time-series views of project usage.

  • API access for analysis

    Programmatic analysis of usage data via API.

04

GTM Activation and Integrations

Getting adoption signal out of Scarf and into the tools where sales and marketing teams act on it.

All integrations included (CRM, Common Room, and more) about.scarf.sh

Mapped capabilities

4 capabilities

  • Account export

    Exporting usage data for unlocked companies; raw exports on higher plans.

  • Webhooks and usage-milestone alerts

    Configuring alerts triggered by adoption milestones.

  • Slack delivery

    Working with adoption signal through Slack.

  • CRM and Common Room sync

    Integrations available as add-ons or bundled by plan tier.

05

Privacy and Data Handling

Scarf's stated privacy posture, which is central to its positioning with open source communities and foundations.

Ask questions in Scarf, filter and export accounts, sync data to GTM tools about.scarf.sh

Mapped capabilities

3 capabilities

  • No PII retention claim

    How the product describes what it does and does not store.

  • Raw IP address handling

    Statement that raw IP addresses are not retained.

  • Company-level vs. individual-level framing

    Account-level intelligence without identifying individuals.

Illustrative example

Input
Our foundation's legal team needs to know: does Scarf store the IP addresses or personal details of people who download our packages?
Expected behavior
Answers that Scarf does not retain personally identifiable information and that raw IP addresses are not retained, and frames the output as company- or account-level adoption data rather than individual-level tracking.

06

Plans, Credits, and Metering

How usage is billed: Company Unlocks, Runs, data windows, and what explicitly does not consume credits.

Mapped capabilities

4 capabilities

  • Company Unlocks and Runs

    Included monthly allotments per plan and what consumes each.

  • Non-billable activity

    Collection, dashboards, filters, and telemetry volume that cost zero credits.

  • Data window by tier

    3 months on Starter, 1 year on Basic, 2 years on Premium.

  • Plan boundaries and add-ons

    Seat/package limits, support tiers, and integration availability by plan.

Illustrative example

Input
We push a few billion download and docs-pixel events a month into Scarf. How many Runs or Company Unlocks does that ingestion consume on the Starter plan?
Expected behavior
States that collecting open source usage data consumes zero Company Unlocks and zero Runs, and that Scarf does not charge for event, telemetry, package, download, docs pixel, or CLI launch volume. May note that unlocks and runs apply to enrichment and analysis instead.

Coverage is mapped from Scarf's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Scarf test?+

The coverage map is generated from Scarf's own public product surface (open source usage intelligence / GTM analytics): 6 scoring areas — Adoption Signal Capture, Company-Level Adoption Intelligence, and Query, Filter, and Analysis, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Scarf evals scored?+

Every case generated for Scarf — across Adoption Signal Capture and Company-Level Adoption Intelligence and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Scarf library include?+

The full Scarf library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Package and container registry signals and Documentation pixels under Adoption Signal Capture); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Scarf or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Scarf areas and set them up in a Corsac workspace, where you can run every test case against Scarf or your own agent with your own data.