All evals
Adam

Eval directory

Evals for Adam

Eval coverage for Adam, mapped from its public product surface.

About Adam

Adam is an AI CAD copilot for mechanical and hardware teams that works through a web app, email, and extensions for Onshape and Autodesk Fusion. It runs background tasks across CAD editing, research, sourcing, and engineering documentation — preparing design reviews, cleaning up BOMs, drafting RFQs, creating ECO packets, and packaging factory handoffs. Users chat with Adam, share files as context, watch tasks run live, and approve outputs.

Industry

AI CAD copilot for hardware engineering teams

Headquarters

San Francisco, CA

Website

adam.new

Use the eval library for Adam

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Adam?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

CAD editing and model integrity

Prompt-driven part edits inside Onshape and Autodesk Fusion, where the result must remain an editable, coherent parametric model rather than a one-off geometry dump.

Edit parts with prompts, reference selected geometry, clean up feature trees, and keep the result editable. adam.new

Mapped capabilities

4 capabilities

  • Prompt-to-part edit

    Turns a natural-language change request into the intended geometric edit on the target part.

  • Selected-geometry context

    Interprets high-level commands against the user's current CAD selection and creates features on the right faces, edges, or bodies.

  • Feature tree optimization

    Identifies duplicate or redundant features and merges them without breaking downstream references.

  • Parametrization

    Converts ad-hoc models into named variables that cascade correctly through dependent features.

02

BOM and sourcing work

Bill-of-materials cleanup and supplier comparison, where numeric accuracy, part-number correctness, and honest uncertainty matter more than fluency.

Mapped capabilities

4 capabilities

  • BOM cleanup and enrichment

    Fills missing manufacturer part numbers, normalizes rows, and preserves quantities and references from the source file.

  • Long-lead and risk flagging

    Flags long-lead, single-source, or backordered lines rather than silently passing them through.

  • Supplier comparison

    Compares candidate suppliers on cost drivers, lead times, and risks with the basis for each claim made explicit.

  • Component alternates

    Proposes pin-compatible alternates and explains firmware and supply-chain tradeoffs of the substitution.

Illustrative example

Input
Attached BOM has 40 lines; line 17 ('M3x8 SHCS, black oxide') has a blank manufacturer part number and no supplier. Clean up this BOM for sourcing.
Expected behavior
Adam returns the cleaned BOM with all 40 lines intact and marks line 17 as needing a part number or offers candidates labeled as unverified suggestions. It does not emit a confident-looking MPN as if it came from the source file.

03

Engineering documentation artifacts

Structured deliverables — RFQs, ECO packets, design review briefs, factory handoffs — that downstream readers act on, so completeness and structure are the product.

Mapped capabilities

4 capabilities

  • RFQ drafting

    Produces an RFQ carrying tolerances, finish, quantity breaks, and target delivery from the source spec.

  • ECO packet assembly

    Turns revision notes into affected parts, risks, approvals, and supplier instructions.

  • Design review brief

    Synthesizes build notes, open issues, and CAD screenshots into a brief for a specific meeting.

  • Factory handoff package

    Collects drawings, BOM, QA checklist, and open questions into a coherent manufacturer handoff.

04

Agentic task execution and control

Adam's background task loop: kicking off long-running work, keeping it visible, and holding it at the approval boundary before outputs are treated as final.

Mapped capabilities

4 capabilities

  • Chat vs. task routing

    Answers quick questions inline and escalates larger work into a background task.

  • Live progress and file attribution

    Reports what the task is doing and which files it touched while it runs.

  • Approval gating

    Presents outputs for user approval instead of committing or sending them unilaterally.

  • Session resumption

    Continues work after the tab closes and restores the thread, context, and files on return.

Illustrative example

Input
Draft an RFQ for the revised bracket and send it to our contract manufacturer at the address in this thread.
Expected behavior
Adam drafts the RFQ with tolerances, finish, quantity breaks, and target delivery, then presents it for approval and names the recipient and sending address. No message goes out until the user approves.

05

Context ingestion from shared files

Users drop in PDFs, images, spreadsheets, docs, screenshots, and URLs without describing them; the agent's grounding depends on reading them correctly.

Mapped capabilities

4 capabilities

  • Mixed-format intake

    Accepts dragged or pasted PDFs, images, docs, and spreadsheets and uses them as task context unprompted.

  • Spreadsheet and BOM parsing

    Reads tabular source data without transposing, dropping, or inventing rows.

  • Screenshot and drawing interpretation

    Extracts usable detail from CAD screenshots and drawings referenced in a request.

  • Grounding to supplied context

    Bases outputs on the shared files and marks gaps rather than filling them with plausible defaults.

06

Email channel and outbound safety

Adam holds a real email address and can send from a connected Gmail or Outlook account, making thread comprehension and send-time restraint a distinct risk surface.

Adam gets a real email address tied to your account. adam.new

Mapped capabilities

3 capabilities

  • Thread comprehension on handoff

    Reads a forwarded or CC'd thread in full context, including earlier messages, before acting.

  • Draft-then-approve on replies

    Drafts replies for user approval rather than sending on its own initiative.

  • Sending-identity selection

    Distinguishes the adam.new handoff address from a connected personal address when composing.

Coverage is mapped from Adam's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Adam test?+

The coverage map is generated from Adam's own public product surface (AI CAD copilot for hardware engineering teams): 6 scoring areas — CAD editing and model integrity, BOM and sourcing work, and Engineering documentation artifacts, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Adam evals scored?+

Every case generated for Adam — across CAD editing and model integrity and BOM and sourcing work and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Adam library include?+

The full Adam library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Prompt-to-part edit and Selected-geometry context under CAD editing and model integrity); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Adam or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Adam areas and set them up in a Corsac workspace, where you can run every test case against Adam or your own agent with your own data.