All evals
Puzzle

Eval directory · Accounting & Finance

Evals for Puzzle

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Puzzle AI products.

About Puzzle

Puzzle is an AI-native accounting platform that automates bookkeeping and financial close for startups and growing companies. Its software ingests transactions, reconciles accounts, and surfaces anomalies in real time — reducing close time from weeks to days.

Employees

~60

Industry

Accounting Software

Headquarters

San Francisco, CA

Website

puzzle.io

Use the eval library for Puzzle

All 161 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Puzzle?

6 areas · 161 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Accounting Close Controls V1

Operational response/safety eval for Puzzle covering accounting close controls.

Mapped capabilities

12 scenarios

  • Close Control Coverage
  • Reconciliation Safety
  • Operator Actionability

Example criterion: Puzzle enforces accounting close controls with reliable exception detection and action-ready finance guidance.

02

Audit Readiness Traceability V1

Wave 2 production eval for Puzzle focused on audit readiness traceability.

Mapped capabilities

16 scenarios

  • Evidence Traceability
  • Control Narrative Quality
  • Risk Exposure Detection

Example criterion: Puzzle maintains audit-ready traceability with defensible evidence links and early detection of control gaps.

03

Ingest Painpoint Eval V1

Evaluates Puzzle's Ingest Painpoint Eval — ingest pipeline fault detection, evidence-linked diagnosis, and remediation prioritization — across 12 test cases graded case by case by an LLM judge.

Mapped capabilities

12 scenarios

  • Ingest Pipeline Fault Detection
  • Evidence-Linked Diagnosis
  • Remediation Prioritization

Example criterion: Puzzle detects ingest painpoints with traceable evidence and prioritized fixes that improve downstream quality.

04

Expert Safety Gate Eval V2 High Conf

Evaluates Puzzle's Expert Safety Gate Eval V2 High Conf — safety gate enforcement, high-risk scenario handling, and release readiness assurance — across 36 test cases graded case by case by an LLM judge.

Mapped capabilities

36 scenarios

  • Safety Gate Enforcement
  • High-Risk Scenario Handling
  • Release Readiness Assurance

Example criterion: Puzzle enforces expert safety gates reliably and prevents unsafe outputs in production-critical scenarios.

05

Power User Ops Eval V2 High Conf

Evaluates Puzzle's Power User Ops Eval V2 High Conf — advanced workflow reliability, safety control integrity, and operational consistency — across 40 test cases graded case by case by an LLM judge.

Mapped capabilities

40 scenarios

  • Advanced Workflow Reliability
  • Safety Control Integrity
  • Operational Consistency

Example criterion: Puzzle supports power users with reliable operations, strong safety controls, and consistent production-quality execution.

06

Workflow Painpoint Eval V2 High Conf

Evaluates Puzzle's Workflow Painpoint Eval V2 High Conf — workflow friction detection, severity prioritization, and actionable fix design — across 45 test cases graded case by case by an LLM judge.

Mapped capabilities

45 scenarios

  • Workflow Friction Detection
  • Severity Prioritization
  • Actionable Fix Design

Example criterion: Puzzle consistently surfaces true workflow painpoints, prioritizes them correctly, and provides remediation steps teams can act on.

Frequently asked questions

What do the Corsac evals for Puzzle test?+

Each eval pack tests Puzzle's public product surface — including Accounting Close Controls V1, Audit Readiness Traceability V1, Ingest Painpoint Eval V1 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Puzzle evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Puzzle library include?+

The Puzzle eval library includes 161 graded test cases across 6 eval packs. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Puzzle or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against Puzzle or your own agent with your own data.