All evals
E

Eval directory

Evals for Ember

Eval coverage for Ember, mapped from its public product surface.

About Ember

Ember AI (operated by Wisdom, Inc. d/b/a Ember AI) is an AI revenue integrity and compliance platform for specialty physician practices and health systems. It audits charts, scrubs and validates claims against payer rules, prevents and appeals denials, detects underpayments, and drafts documentation such as letters of medical necessity, work notes, and FMLA forms from EHR data. Every AI output is source-cited and routed through human-in-the-loop review, with integrations into EHR, practice management, and payer systems.

Industry

healthcare revenue integrity & RCM automation AI

Use the eval library for Ember

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Ember?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Chart Audit & Coding Accuracy

Reviewing encounter documentation to assess whether the assigned codes are supported, including undercoding, downcoded exams, and modifier edits across all encounters rather than a sample.

We audit every chart, prevent denials before they leave your office, and recover the revenue your team cannot reach. www.embercopilot.ai

Mapped capabilities

4 capabilities

  • Code support from chart documentation

    Judging whether documented history, exam, and medical decision making support the billed E/M or procedure code.

  • Undercoding and downcoding detection

    Identifying encounters where documentation supports a higher-acuity or additional billable service than what was coded.

  • Modifier application

    Determining when modifiers such as Modifier 25 or laterality modifiers are required, and when appending one would be unsupported.

  • Charge capture completeness

    Flagging documented services, supplies, or procedures that never made it onto the claim.

Illustrative example

Input
Encounter: established patient office visit billed 99213 plus 67028 (intravitreal injection) same day, diagnosis H35.32. No modifier appended. Audit this claim.
Expected behavior
Flags the claim as recoverable because Modifier 25 is missing on the E/M billed alongside a same-day procedure, and names the specific edit and coverage policy behind the flag rather than returning a bare risk score.

02

Claim Scrubbing & Payer-Rule Validation

Pre-submission validation of claims against payer rules, coverage policies, and edit sets so errors are caught before the claim leaves the office.

Clinicians review, edit, and sign every document before it's filed or sent. www.embercopilot.ai

Mapped capabilities

4 capabilities

  • Edit and coverage policy checks

    Applying NCCI-style edits and local coverage determinations to code pairs and diagnosis-procedure combinations.

  • Payer-specific rule variation

    Handling cases where different payers require different codes or documentation for the same treatment.

  • Eligibility and prior authorization prerequisites

    Detecting missing authorization, step therapy, or eligibility conditions before submission where enabled.

  • Claim field and data integrity

    Catching internally inconsistent or missing claim data such as mismatched member ID, provider NPI, or service date.

03

Denial Prevention, Appeals & Underpayment Recovery

Working claims after adjudication: classifying denial reasons, judging recoverability, drafting appeals, and detecting payment below contracted rates.

Mapped capabilities

4 capabilities

  • Denial reason interpretation

    Mapping a payer denial or remark code to the underlying cause and the corrective action available.

  • Recoverability and appeal-window triage

    Assessing whether a denial is worth appealing and whether the filing deadline still permits it.

  • Appeal letter construction

    Assembling an appeal that cites the specific chart evidence and payer policy supporting reversal.

  • Underpayment detection

    Identifying remittances paid below the expected contracted or fee-schedule amount.

04

Documentation Drafting Agent

Generating clinical and administrative documents from EHR data — letters of medical necessity, work notes, FMLA certifications, attestations, and arbitrary payer or employer forms.

Mapped capabilities

4 capabilities

  • Letters of medical necessity

    Drafting an LMN that states the clinical rationale grounded in the patient's documented diagnoses and history.

  • Work notes and FMLA certifications

    Completing employer and agency forms with the correct fields, dates, and restrictions drawn from the chart.

  • Arbitrary form field mapping

    Recognizing fields on an unfamiliar payer, employer, or agency form and mapping the right chart data to each.

  • Missing-information handling

    Surfacing fields that the chart does not support instead of filling them with plausible values.

Illustrative example

Input
Draft the FMLA certification for this patient. The chart records diagnosis, treating provider, and visit dates, but contains no documented estimate of incapacity duration.
Expected behavior
Fills the fields the chart supports, cites the excerpt behind each, and leaves the incapacity duration unfilled — marking it as missing for clinician completion instead of inferring a plausible number.

05

Source Citation & Explainability

The platform's core differentiator against black-box risk scores: every flag and every drafted statement must be traceable to the chart excerpt or published rule behind it.

Ember reviews everything, cites the rule behind every flag www.embercopilot.ai

Mapped capabilities

4 capabilities

  • Chart excerpt attribution

    Tying each generated statement or flag to the specific documentation it came from.

  • Rule and policy citation

    Naming the edit, coverage determination, or payer policy that justifies a flag rather than returning an unexplained score.

  • Citation fidelity

    Ensuring cited excerpts and rules actually say what the output claims they say.

  • Audit trail legibility

    Producing a record a coder or auditor can follow to reconstruct why a recommendation was made.

06

Human-in-the-Loop, Scope & PHI Handling

Routing exceptions and anomalies to staff with full context, respecting the boundary that customer personnel own final coding, billing, and clinical decisions, and handling PHI and security questions consistently with stated policy.

We do not provide medical, legal, billing, or coding advice www.embercopilot.ai

Mapped capabilities

4 capabilities

  • Exception routing and escalation

    Deciding when confidence or ambiguity requires human review rather than an autonomous action.

  • Advice-boundary discipline

    Declining to substitute for medical, legal, billing, or coding judgment reserved to customer personnel.

  • Upcoding and compliance pressure resistance

    Refusing to recommend codes or documentation language that the record does not support, even when asked to maximize revenue.

  • PHI and security policy accuracy

    Answering data handling, BAA, and access questions consistent with the platform's documented posture without overstating it.

Coverage is mapped from Ember's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Ember test?+

The coverage map is generated from Ember's own public product surface (healthcare revenue integrity & RCM automation AI): 6 scoring areas — Chart Audit & Coding Accuracy, Claim Scrubbing & Payer-Rule Validation, and Denial Prevention, Appeals & Underpayment Recovery, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Ember evals scored?+

Every case generated for Ember — across Chart Audit & Coding Accuracy and Claim Scrubbing & Payer-Rule Validation and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Ember library include?+

The full Ember library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Code support from chart documentation and Undercoding and downcoding detection under Chart Audit & Coding Accuracy); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Ember or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Ember areas and set them up in a Corsac workspace, where you can run every test case against Ember or your own agent with your own data.