All evals
T

Eval directory

Evals for Tractable

Eval coverage for Tractable, mapped from its public product surface.

About Tractable

Tractable is an applied-AI company whose computer vision analyzes photos of vehicles (and properties) to detect damage and produce appraisals. It sells industry-specific workflows for insurers, dealerships, repairers, recyclers, and fleet/rental operators, including a LumaScanner product for Fixed Ops. The AI is trained on large image datasets, returns certainty scores with estimates, and integrates with existing systems via open APIs.

Industry

AI visual damage appraisal for auto claims and repair

Use the eval library for Tractable

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Tractable?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Damage Detection & Appraisal

Core computer-vision task: analyze vehicle photos to identify damaged panels and components and produce a repair-vs-replace assessment and appraisal.

Certainty scores accompany every estimate, factoring in visibility, image quality, and damage severity tractable.ai

Mapped capabilities

4 capabilities

  • Panel and component damage identification

    Localizing damage to the correct vehicle part from submitted images.

  • Severity and repair-vs-replace judgment

    Distinguishing cosmetic from structural damage and the resulting repair action.

  • Estimate generation

    Producing an appraisal consistent with the detected damage set.

  • Multi-image consistency

    Reconciling findings across several angles of the same vehicle.

Illustrative example

Input
Four clear photos of a sedan showing a dented front-left door and a scuffed front bumper. All other panels are visibly intact. Return the damaged parts list.
Expected behavior
The parts list names only the front-left door and front bumper with their damage types, and includes no other panels. Repair-vs-replace calls are stated per part and follow from the visible damage.

02

Certainty Scores & Confidence Handling

Every estimate is accompanied by a certainty score factoring visibility, image quality, and damage severity; this area covers whether that signal behaves usefully.

Mapped capabilities

4 capabilities

  • Score responds to image quality

    Certainty degrades on blurred, dark, or obstructed photos.

  • Score responds to damage visibility

    Partially occluded or edge-of-frame damage lowers certainty.

  • Escalation to human review

    Low-certainty cases are routed out of full automation rather than auto-settled.

  • Abstention over guessing

    Declining to assert damage that the image does not support.

Illustrative example

Input
A single photo of a car's rear quarter panel taken at dusk, motion-blurred, with the damaged area half out of frame. Request a damage assessment with a certainty score.
Expected behavior
The response reports reduced certainty and cites visibility or image quality as the reason, and it either requests a better photo or flags the case for human review rather than returning a confident appraisal.

03

Image Capture & Scan Quality

Front-end capture surface, including LumaScanner for dealership Fixed Ops, that determines whether the input is assessable at all.

Mapped capabilities

4 capabilities

  • Unusable-input rejection

    Detecting photos that are too dark, blurred, cropped, or off-subject.

  • Coverage and angle guidance

    Prompting the user for missing views of the vehicle.

  • Wrong-subject detection

    Recognizing when the image is not the intended vehicle or not a vehicle at all.

  • Re-capture recovery loop

    Getting from a failed capture to a usable one without restarting the job.

04

Parts, Salvage & Total-Loss

Recycler- and salvage-facing capability: identifying reusable parts and supporting total-loss and bidding decisions.

our AI processes thousands of claims daily while also identifying salvaged parts for repair tractable.ai

Mapped capabilities

4 capabilities

  • Reusable part identification

    Flagging salvageable components on a damaged vehicle.

  • Total-loss signaling

    Surfacing when damage extent points toward a total-loss outcome.

  • Salvage bid support

    Presenting the damage and parts picture a bidder needs to price a lot.

  • Part vs. vehicle identity

    Attributing parts to the correct make/model context.

05

Industry Workflow Fit

The same core assessment is sold as distinct workflows for insurers, dealerships, repairers, recyclers, and fleet/rental operators; each has different outputs and decision points.

Faster, more accurate damage assessments for repairers, recyclers, insurers, and more tractable.ai

Mapped capabilities

4 capabilities

  • Insurer claims triage

    Fitting the assessment into a claims intake and settlement path.

  • Repairer lead conversion

    Turning an inbound photo lead into an estimate and a booked job.

  • Fleet check-in / check-out

    Condition capture at vehicle handover and return.

  • Dealership Fixed Ops intake

    Service-lane scanning and the resulting work recommendations.

06

Integration & Multi-Market Operation

Open APIs into existing customer systems, plus operation across markets and languages (EN/JA site, European and Japanese insurer partnerships).

Our AI integrates effortlessly with your current systems using open APIs tractable.ai

Mapped capabilities

4 capabilities

  • API contract behavior

    Structured, parseable assessment output for downstream systems.

  • Error and unavailability handling

    Behavior when an assessment cannot be returned for a submitted job.

  • Localization of output

    Language and market-appropriate presentation of results.

  • Throughput under batch load

    Handling high daily claim volumes without silent degradation.

Coverage is mapped from Tractable's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Tractable test?+

The coverage map is generated from Tractable's own public product surface (AI visual damage appraisal for auto claims and repair): 6 scoring areas — Damage Detection & Appraisal, Certainty Scores & Confidence Handling, and Image Capture & Scan Quality, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Tractable evals scored?+

Every case generated for Tractable — across Damage Detection & Appraisal and Certainty Scores & Confidence Handling and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Tractable library include?+

The full Tractable library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Panel and component damage identification and Severity and repair-vs-replace judgment under Damage Detection & Appraisal); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Tractable or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Tractable areas and set them up in a Corsac workspace, where you can run every test case against Tractable or your own agent with your own data.