All evals
D

Eval directory

Evals for Distro

Eval coverage for Distro, mapped from its public product surface.

About Distro

Distro is an AI revenue platform for wholesale distributors that automates quoting, product lookup, and knowledge retrieval across counter, inside, outside, and commercial sales teams. It ingests RFQs, plans, and spec packages in any format and turns them into matched, priced quotes and structured BOMs, with CPQ questionnaires and built-in engineering logic for technical products. It connects to a distributor's ERP and product data during a guided implementation, and ships pre-loaded with manufacturer technical documentation.

Industry

AI sales automation and quoting for wholesale distributors

Website

distro.app

Use the eval library for Distro

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Distro?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

RFQ intake and quote generation

Turning inbound requests-for-quote of any format and length into a completed, priced quote, including parsing line items, quantities, and units out of unstructured text and attachments.

Turn a 100+ line RFQ into a completed quote in 90 seconds. distro.app

Mapped capabilities

4 capabilities

  • Multi-format RFQ parsing

    Extract line items, quantities, and units of measure from emails, PDFs, spreadsheets, and photos of handwritten lists without dropping or duplicating lines.

  • Long multi-line RFQ handling

    Preserve completeness and ordering across RFQs of 100+ lines, including repeated items, subtotals, and continuation pages.

  • Ambiguous or incomplete line resolution

    Flag lines that lack enough detail to quote rather than guessing a part, and state what information is missing.

  • Quote assembly and output structure

    Produce a quote whose line items, quantities, and totals reconcile to the source request.

02

Product matching and cross-reference

Mapping a requested item — by description, competitor part number, or manufacturer number — to the correct item in the distributor's catalog, with an honest confidence signal.

Mapped capabilities

4 capabilities

  • Manufacturer part number matching

    Resolve exact, partial, and mis-typed manufacturer numbers to the right catalog SKU.

  • Competitor cross-referencing

    Identify a valid equivalent when the requested brand is not carried, and distinguish exact equivalents from approximate substitutes.

  • Description-only matching

    Match free-text descriptions with sizes, materials, and ratings to catalog items when no part number is supplied.

  • Confidence scoring and abstention

    Surface low-confidence matches for human review instead of committing to a wrong item.

Illustrative example

Input
Customer email: "Need 12 of Acme part CR-4402-B, whatever you carry that crosses to it." The catalog carries no exact equivalent, only a close alternative with a different pressure rating.
Expected behavior
The response offers the close alternative as an approximate substitute, explicitly names the pressure-rating difference, and does not present it as an exact cross-reference. Confidence is marked below the threshold for automatic quoting so a rep reviews the line.

03

Pricing and catalog data fidelity

Applying the distributor's own pricing, availability, and customer-specific terms from connected ERP and product data rather than inventing values.

Mapped capabilities

4 capabilities

  • Customer-specific pricing

    Apply the correct price for the identified customer, contract, or price level.

  • Cost change and repricing behavior

    Reflect updated costs when quoting, and avoid quoting from stale figures without saying so.

  • Availability and substitution signaling

    State stock or lead-time position where the connected data supports it, and avoid asserting it where it does not.

  • No-fabrication discipline

    Decline to state a price or spec that is not present in the connected catalog or documentation.

04

Plan and spec to structured BOM

Reading multi-page specification books, equipment schedules, and drawings and producing a clean, priced material list matched to the catalog.

From messy plans and specs to clean, structured BOMs in minutes. distro.app

Mapped capabilities

4 capabilities

  • Schedule and spec extraction

    Pull equipment tags, models, capacities, and quantities out of schedules and specification sections.

  • Spec compliance in matching

    Honor stated specification constraints when selecting catalog items rather than matching on model number alone.

  • Coverage and omission reporting

    Report which portions of the package were priced and which were not, so nothing is silently skipped.

  • BOM structure and traceability

    Return a structured material list where each line traces back to its source page or schedule.

Illustrative example

Input
A 47-page mechanical spec package where one equipment schedule page is a skewed, low-resolution scan with illegible model numbers. The rest of the package is clean.
Expected behavior
The BOM prices every legible line and reports the unreadable page as unpriced, naming the page and what could not be read. It does not infer model numbers from surrounding context or silently omit the page from the coverage summary.

05

Technical CPQ and engineering logic

Guiding a rep through structured, category-specific questionnaires and applying built-in engineering logic for sizing, compatibility, and configuration of technical products.

Built-in logic handles the engineering: load calculations, equipment sizing, component compatibility. distro.app

Mapped capabilities

4 capabilities

  • Questionnaire sequencing

    Ask the questions that actually narrow the configuration for the product category at hand, and stop when enough is known.

  • Sizing and load calculation

    Apply the correct calculation for the stated application inputs and show the basis for the result.

  • Compatibility and accessory completeness

    Include required components and reject combinations that do not work together.

  • Insufficient-input handling

    Refuse to size or configure when a required input is missing, and name the missing input.

06

Knowledge retrieval and technical answers

Answering product and technical questions for counter, phone, and field use from pre-loaded manufacturer documentation and enriched product data, with sourcing a rep can defend to a customer.

80%+ Faster Knowledge Retrieval distro.app

Mapped capabilities

4 capabilities

  • Spec lookup from manufacturer docs

    Return the correct value from the applicable manufacturer document for the product in question.

  • Citation and source attribution

    Point to the document and location that supports the answer.

  • Out-of-scope and unsupported questions

    Say when the answer is not in the available documentation instead of improvising.

  • Conflicting or superseded documentation

    Prefer the applicable revision and surface the conflict when sources disagree.

Coverage is mapped from Distro's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Distro test?+

The coverage map is generated from Distro's own public product surface (AI sales automation and quoting for wholesale distributors): 6 scoring areas — RFQ intake and quote generation, Product matching and cross-reference, and Pricing and catalog data fidelity, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Distro evals scored?+

Every case generated for Distro — across RFQ intake and quote generation and Product matching and cross-reference and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Distro library include?+

The full Distro library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Multi-format RFQ parsing and Long multi-line RFQ handling under RFQ intake and quote generation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Distro or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Distro areas and set them up in a Corsac workspace, where you can run every test case against Distro or your own agent with your own data.