All evals
C

Eval directory

Evals for CodaMetrix

Eval coverage for CodaMetrix, mapped from its public product surface.

About CodaMetrix

CodaMetrix is an AI-powered contextual coding automation platform for health systems, marketed under the CMX CARE™ name. It converts clinical documentation into a longitudinal, patient-centric view and automatically applies billable codes across facility-based and professional fee service lines. Founded in 2019 out of Mass General Brigham physician organizations, it targets lower coding cost, faster turnaround, and fewer coding-related denials.

Industry

AI medical coding automation for healthcare revenue cycle

Use the eval library for CodaMetrix

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for CodaMetrix?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Contextual Code Assignment

Converting clinical documentation into billable codes across facility-based and professional fee service lines, using a longitudinal, patient-centric view rather than a single-document read.

Automatically applies codes across service lines—supporting care, compliance, and performance—without added workflows. www.codametrix.com

Mapped capabilities

4 capabilities

  • Professional fee coding

    Assigns correct procedure and diagnosis codes from physician documentation for specialty service lines such as radiology and pathology.

  • Facility-based coding

    Applies facility-side codes from encounter documentation, including the emergency department solution surface.

  • Longitudinal context synthesis

    Uses prior and concurrent documentation for the same patient to code correctly when a single note is insufficient.

  • Coding at first opportunity

    Produces a code as soon as the relevant clinical data is present rather than waiting for a complete downstream record.

Illustrative example

Input
A follow-up CT report reads 'stable, no interval change' with no stated indication. The patient's prior imaging report in the record documents a known pulmonary nodule.
Expected behavior
The system uses the prior report to code the known pulmonary nodule as the diagnosis rather than defaulting to an unspecified or screening code, and cites the prior report as the supporting evidence.

02

Coding Quality and Compliance

Producing codes that hold up to audit: supported by documentation, consistent with coding guidelines, and explainable at the code level.

Mapped capabilities

4 capabilities

  • Documentation-supported codes

    Refuses to assign codes that the underlying documentation does not substantiate.

  • Guideline and specificity adherence

    Selects the most specific defensible code and respects code-set and modifier conventions.

  • Code-level rationale

    Cites the documentation evidence behind each assigned code for auditor review.

  • Consistency across encounters

    Applies the same coding decision to equivalent documentation rather than varying by run or service line.

03

Denial Prevention and Turnaround

Surfacing the documentation and coding problems that drive coding-related denials before claim submission, and delivering codes fast enough to shorten revenue cycle turnaround.

CodaMetrix delivers the power of high-performance medical coding—up to 30% lower cost. www.codametrix.com

Mapped capabilities

4 capabilities

  • Denial-risk detection

    Flags code and documentation combinations likely to be denied by payers.

  • Medical necessity linkage

    Checks that diagnosis support for the billed service is present in the documentation.

  • Charge capture completeness

    Identifies chargeable activity present in documentation but missing from the coded output.

  • Turnaround behavior

    Handles late-arriving or amended documentation without stalling or silently reworking finalized codes.

04

Coder Workflow and Escalation

The division of labor between automation and the health system's coding team: what is coded autonomously, what is routed to a human, and how coder decisions feed back into the system.

Mapped capabilities

4 capabilities

  • Autonomy thresholds

    Codes autonomously only where confidence and evidence support it, and escalates otherwise.

  • Escalation packaging

    Hands a coder the specific ambiguity and supporting documentation rather than an unexplained rejection.

  • Learning from coder corrections

    Incorporates coding decisions made by the health system's team into subsequent behavior.

  • Backlog and queue handling

    Prioritizes and drains coding backlogs across specialties without dropping work.

05

Clinical Data Ingestion and EMR Fit

Taking in documentation from the health system's existing EMR and workflows with minimal technical lift, and resolving it to the right patient and encounter.

CodaMetrix Chosen By Health Systems Representing $180B In Net Patient Revenue www.codametrix.com

Mapped capabilities

4 capabilities

  • Multi-EMR document intake

    Ingests documentation from major EMRs such as Epic, Cerner, Athena, Allscripts, and Meditech.

  • Patient and encounter resolution

    Attaches documents to the correct patient and encounter when building the longitudinal view.

  • Workflow fit

    Delivers coded output into existing coding workflows without requiring new manual steps.

  • Service line scaling

    Extends to additional specialties and service lines without per-specialty rework.

06

Failure Handling and Safe Fallback

How the platform behaves when documentation is incomplete, contradictory, or unavailable — the cases where a wrong autonomous code creates compliance exposure.

Mapped capabilities

4 capabilities

  • Abstention on insufficient evidence

    Declines to code and routes to a human when documentation does not support any confident assignment.

  • Contradictory documentation

    Detects conflicts between documents about the same encounter instead of silently choosing one.

  • Missing or truncated documents

    Recognizes an incomplete record and withholds a final code rather than guessing.

  • No inference beyond the record

    Avoids asserting clinical findings or services not present in the supplied documentation.

Illustrative example

Input
An operative note ends mid-sentence during the procedure description, so the approach and any additional procedures performed are not documented anywhere in the record.
Expected behavior
The system does not finalize a procedure code. It flags the record as incomplete and escalates to a human coder, naming the missing approach and procedure detail as the reason.

Coverage is mapped from CodaMetrix's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for CodaMetrix test?+

The coverage map is generated from CodaMetrix's own public product surface (AI medical coding automation for healthcare revenue cycle): 6 scoring areas — Contextual Code Assignment, Coding Quality and Compliance, and Denial Prevention and Turnaround, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the CodaMetrix evals scored?+

Every case generated for CodaMetrix — across Contextual Code Assignment and Coding Quality and Compliance and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the CodaMetrix library include?+

The full CodaMetrix library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Professional fee coding and Facility-based coding under Contextual Code Assignment); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against CodaMetrix or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped CodaMetrix areas and set them up in a Corsac workspace, where you can run every test case against CodaMetrix or your own agent with your own data.