All evals
F

Eval directory

Evals for Fathom

Eval coverage for Fathom, mapped from its public product surface.

About Fathom

Fathom is an AI system that automates medical coding for health systems, physician groups, clinics, health plans, value-based care providers, and RCM vendors. It offers direct-to-bill coding automation that processes charts without human intervention and routes the remainder to human coders, plus an 'Audit complete' service that reviews coded charts for denial and downcoding risk. The company markets it on cost reduction, speed, scale, and accuracy, and cites HITRUST, HIPAA, and SOC 2 Type 2 compliance.

Industry

autonomous medical coding AI

Headquarters

Offices listed in Oakland, CA and New York, NY; office hubs in New York, San Francisco, and Toronto (no page labels a headquarters)

Use the eval library for Fathom

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Fathom?

6 scoring areas · 22 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Autonomous Coding Accuracy

Assigning correct codes directly from chart documentation for professional and facility coding across service lines, the core of the direct-to-bill offering.

Direct to bill. No human intervention required. www.fathomhealth.com

Mapped capabilities

4 capabilities

  • E/M level assignment

    Selects the evaluation and management level supported by documented history, complexity, and time, without upcoding or downcoding.

  • Provider assignment

    Attributes the encounter to the correct rendering and supervising provider so clinicians are paid accurately.

  • Diagnosis and procedure code selection

    Emits ICD and procedure codes that the chart text actually supports, at the correct specificity.

  • Facility versus professional context

    Applies the coding rules appropriate to the setting rather than mixing facility and professional conventions.

Illustrative example

Input
Established patient office visit: focused history, one stable chronic condition reviewed, no medication changes, no new orders, 12 minutes total time documented. Assign the professional E/M code.
Expected behavior
Assigns the low-complexity established-patient E/M code supported by the documented complexity and time, and does not select a higher level absent supporting documentation.

02

Automation Routing and Human Handoff

The decision of which charts are coded end-to-end without human intervention and which are passed to the client's existing coding operation.

Mapped capabilities

4 capabilities

  • Direct-to-bill eligibility

    Codes only charts whose documentation supports an unattended decision and releases them to billing.

  • Routing to human coders

    Passes ambiguous, incomplete, or atypical charts to the coding team instead of guessing.

  • Abstention on insufficient documentation

    Declines to assign a code when the chart lacks the supporting detail, rather than inferring one.

  • Handoff payload quality

    Gives the receiving coder the chart context and reason for routing needed to work the exception.

03

Risk-Adjustment and HCC Capture

Retrospective risk-adjustment coding for health plans and value-based care providers, where exhaustive and timely ICD capture drives RAF accuracy.

Mapped capabilities

4 capabilities

  • Exhaustive HCC capture

    Identifies every HCC-mapping condition the documentation supports across the chart.

  • Unsupported capture suppression

    Does not capture conditions that lack documentation, avoiding inflated RAF scores.

  • ICD specificity

    Selects the most specific ICD code the record supports rather than an unspecified fallback.

  • Chart-level reconciliation

    Keeps captured conditions traceable to the encounter and date of service that supports them.

04

Audit Complete: Denial and Downcoding Review

Reviewing already-coded charts from a coding team or vendor and flagging those that represent denial risk or unnecessary downcoding.

The world’s first comprehensive, real-time coding audit www.fathomhealth.com

Mapped capabilities

4 capabilities

  • Denial-risk flagging

    Flags coded charts whose code set is likely to draw a payer denial.

  • Downcoding detection

    Identifies charts coded below the level the documentation supports.

  • Coding error identification

    Surfaces incorrect or incomplete codes in submitted work and names the specific code at issue.

  • Pass-through discipline

    Leaves correctly coded charts unflagged so reviewers are not buried in false positives.

Illustrative example

Input
Coded chart for audit: the documentation supports a higher-specificity diagnosis than the code the human coder assigned. Review the chart and its submitted code set.
Expected behavior
Flags the chart as downcoding risk, names the submitted code and the documentation supporting the more specific code, and does not return a clean pass verdict.

05

PHI Handling and Compliance Posture

Safeguarding protected health information in processing and outputs, consistent with the HITRUST, HIPAA, and SOC 2 Type 2 posture the company publishes.

verify Fathom's HITRUST, HIPAA and SOC 2 Type 2 compliance www.fathomhealth.com

Mapped capabilities

3 capabilities

  • PHI minimization in outputs

    Includes only the patient information the coding or audit result requires.

  • Auditability of coding decisions

    Ties each assigned or flagged code back to the documentation that supports it for later review.

  • Compliance claim accuracy

    Describes its own certifications and safeguards accurately without overstating scope.

06

Coding Operations Scale and Continuity

Sustaining high-volume throughput and predictable turnaround for coding operations, including behavior when volume spikes or upstream inputs degrade.

Reduce the total cost of your coding operations by up to 50%. www.fathomhealth.com

Mapped capabilities

3 capabilities

  • Throughput under volume

    Maintains coding quality as chart volume scales rather than degrading output.

  • Turnaround consistency

    Returns coded charts within the expected cycle so downstream billing is not delayed.

  • Degraded-input recovery

    Handles malformed, truncated, or duplicate charts without stalling the queue or emitting bad codes.

Coverage is mapped from Fathom's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Fathom test?+

The coverage map is generated from Fathom's own public product surface (autonomous medical coding AI): 6 scoring areas — Autonomous Coding Accuracy, Automation Routing and Human Handoff, and Risk-Adjustment and HCC Capture, and more — spanning 22 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Fathom evals scored?+

Every case generated for Fathom — across Autonomous Coding Accuracy and Automation Routing and Human Handoff and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Fathom library include?+

The full Fathom library is built on request. The coverage map spans 6 areas and 22 capabilities (for example, E/M level assignment and Provider assignment under Autonomous Coding Accuracy); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Fathom or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Fathom areas and set them up in a Corsac workspace, where you can run every test case against Fathom or your own agent with your own data.