All evals
NY

Eval directory

Evals for Nym

Eval coverage for Nym, mapped from its public product surface.

About Nym

Nym is an autonomous medical coding engine that reads clinical language in patient charts and assigns medical codes without human intervention. It runs in the background of existing revenue cycle workflows, taking a first pass at coding encounters and routing successfully coded ones to billing. It uses Clinical Language Understanding (CLU) plus a rules-based approach to code assignment, producing audit trails for each assigned code.

Industry

autonomous medical coding AI (healthcare revenue cycle)

Headquarters

NYC (HQ), with R&D office in Tel Aviv

Website

nym.health

Use the eval library for Nym

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Nym?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Clinical Language Understanding & Code Assignment

The core engine behavior: deciphering narrative clinical language in a chart and assigning the correct diagnosis and procedure codes without human intervention.

accurately assigns medical codes in seconds, all with zero human intervention nym.health

Mapped capabilities

4 capabilities

  • Diagnosis code abstraction from narrative

    ICD-10-CM assignment from free-text documentation, including specificity and laterality.

  • Procedure and service code assignment

    CPT/HCPCS and ICD-10-PCS selection from documented procedures.

  • Modifier and E/M level determination

    CPT modifiers and medical-decision-making-driven E/M leveling.

  • Negation, uncertainty, and historical context

    Ruled-out, suspected, family-history, and resolved conditions distinguished from active ones.

Illustrative example

Input
ED encounter note: patient presents with chest pain; physician documents "likely early pneumonia, will treat empirically." No confirmatory imaging or final diagnosis recorded. Assign ICD-10-CM diagnosis codes.
Expected behavior
The engine codes the documented sign, chest pain, and does not assign a pneumonia code, because outpatient and ED coding guidelines prohibit coding uncertain diagnoses as confirmed. The audit trail names the uncertain-diagnosis rule as the reason.

02

Audit Trail & Explainability

Nym advertises fully transparent audit trails for every code assigned; this area covers whether each code is traceable back to chart evidence and to the rule that produced it.

produces fully transparent audit trails for every code assigned nym.health

Mapped capabilities

4 capabilities

  • Per-code documentation citation

    Each assigned code points to the specific chart language supporting it.

  • Rule traceability

    The coding guideline or rule applied is identified alongside the code.

  • Reviewer-readable rationale

    Explanations a coder or auditor can act on without engineering help.

  • Rationale–code consistency

    Stated justification actually matches the code emitted.

03

Coding Guideline Currency & Compliance

Nym claims automatic updates as soon as new coding guidelines are released. This area covers adherence to current guidelines, edits, and compliance boundaries.

automatically updated as soon as new coding guidelines are released nym.health

Mapped capabilities

4 capabilities

  • Current-year code set adherence

    Retired, revised, and newly added codes handled per the effective guideline year.

  • NCCI edits and bundling

    Bundled, mutually exclusive, and unbundling-prohibited code pairs.

  • Medical necessity and coverage alignment

    Diagnosis-to-procedure support consistent with payer coverage rules.

  • Upcoding and downcoding resistance

    Code level stays anchored to documentation rather than to revenue.

04

Autonomous Routing & Human Handoff

The engine takes a first pass and routes successfully coded encounters to billing. This area covers when it should decline to finalize and hand the encounter back instead.

The engine runs in the background, taking a first pass at coding patient encounters nym.health

Mapped capabilities

4 capabilities

  • Auto-bill versus hold decision

    Only fully and defensibly coded encounters proceed to billing.

  • Abstention on insufficient documentation

    No guessed code when the chart lacks a required element.

  • Exception reason codes for coders

    Held encounters carry an actionable reason for the human reviewer.

  • CDI query candidates

    Documentation gaps surfaced as query-worthy rather than silently coded.

Illustrative example

Input
Operative note documents a knee arthroscopy with meniscectomy but never states left or right. All other required elements are present. Code the encounter and route it.
Expected behavior
The engine withholds the laterality-dependent code and routes the encounter to human review rather than defaulting to a side or emitting an unspecified-laterality code. The exception reason identifies the missing laterality documentation.

05

Workflow Integration & Throughput

Nym runs in the background of existing revenue cycle workflows without disrupting them; this area covers ingestion, handoff, and behavior under production volume.

brings coding automation into the existing workflow without disrupting it nym.health

Mapped capabilities

4 capabilities

  • Chart and encounter ingestion

    Encounter intake from EHR-sourced documentation of varying completeness.

  • Billing system handoff

    Output structure and completeness of the coded encounter passed downstream.

  • Non-disruptive background operation

    First-pass coding coexists with the existing coder workflow.

  • Time-to-bill behavior at volume

    Turnaround and stability across high encounter throughput.

06

Setting & Specialty Coverage

Nym markets multispecialty autonomous coding across health systems and physician groups, with emergency department coding as its flagship case; this area covers breadth and setting-specific rules.

Trusted by 30+ health systems & physician groups nym.health

Mapped capabilities

4 capabilities

  • Emergency department coding

    ED facility and professional coding conventions.

  • Ambulatory surgery and outpatient coding

    Outpatient and ASC coding rules, including OPPS/APC context.

  • Facility versus professional fee coding

    Correct application of facility versus profee conventions for the same encounter.

  • Multispecialty breadth

    Specialty-specific conventions across areas such as radiology, cardiology, and orthopedics.

Coverage is mapped from Nym's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Nym test?+

The coverage map is generated from Nym's own public product surface (autonomous medical coding AI (healthcare revenue cycle)): 6 scoring areas — Clinical Language Understanding & Code Assignment, Audit Trail & Explainability, and Coding Guideline Currency & Compliance, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Nym evals scored?+

Every case generated for Nym — across Clinical Language Understanding & Code Assignment and Audit Trail & Explainability and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Nym library include?+

The full Nym library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Diagnosis code abstraction from narrative and Procedure and service code assignment under Clinical Language Understanding & Code Assignment); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Nym or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Nym areas and set them up in a Corsac workspace, where you can run every test case against Nym or your own agent with your own data.