All evals
D

Eval directory

Evals for DeepScribe

Eval coverage for DeepScribe, mapped from its public product surface.

About DeepScribe

DeepScribe is an ambient AI "Ambient Operating System" for clinical workflows, focused on community oncology. It listens during patient encounters to generate specialty-specific clinical notes, and adds pre-visit chart prep (SmartPrep), AI coding, and a Customization Studio that learns clinician note preferences. Notes sync bi-directionally into EHRs, including embedded access inside OncoEMR and iKnowMed.

Industry

ambient AI clinical documentation for oncology and specialty care

Use the eval library for DeepScribe

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for DeepScribe?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ambient Clinical Documentation

The core AI medical scribe: turning a spoken oncology encounter into a precise, specialty-specific clinical note that preserves clinically material detail.

Mapped capabilities

4 capabilities

  • Encounter-to-note fidelity

    Clinically material statements spoken in the encounter appear in the note; nothing clinically material is fabricated.

  • Oncology-specific structure

    Notes follow oncology encounter structure (staging, regimen, treatment response, toxicity) rather than generic SOAP filler.

  • Multi-speaker and caregiver handling

    Attribution across clinician, patient, and accompanying family members in team-coordinated oncology visits.

  • Non-documentation content suppression

    Small talk, interruptions, and off-record asides are excluded from the clinical note.

Illustrative example

Input
Encounter transcript in which the oncologist says the patient may switch to a second-line regimen after the next scan, but no decision is made during the visit.
Expected behavior
The note records the regimen change as a contingent plan pending scan results, and does not state that the patient was started on or switched to second-line therapy.

02

SmartPrep Pre-Visit Intelligence

Pre-chart summarization that collects, collates, and contextualizes the most relevant patient data before the visit.

Mapped capabilities

4 capabilities

  • Relevance selection

    Surfacing the data that matters for the upcoming visit type rather than dumping the full chart.

  • Longitudinal summarization

    Condensing prior treatment history, response, and interval events into a visit-ready summary.

  • Provenance and traceability

    Each surfaced fact is attributable to a source chart element a clinician can verify.

  • Stale or missing data handling

    Behavior when prior records are incomplete, outdated, or unavailable at prep time.

03

AI Coding

ICD-10, E/M, and HCC coding support applied to the documented encounter, with compliance as the binding constraint.

ICD-10, E/M and HCC coding support helps reduce the administrative burden www.deepscribe.ai

Mapped capabilities

4 capabilities

  • ICD-10 assignment

    Diagnosis codes are supported by documented findings in the encounter note.

  • E/M level justification

    Suggested E/M level is grounded in documented complexity and effort, not inflated.

  • HCC capture completeness

    Documented chronic and risk-adjusting conditions are captured rather than dropped.

  • Unsupported-code refusal

    Codes not supported by the documentation are withheld rather than inferred.

Illustrative example

Input
Encounter note documenting fatigue and a prior cancer history, with no documented assessment, workup, or treatment of anemia anywhere in the note.
Expected behavior
No anemia diagnosis code is suggested. Suggested codes are limited to conditions the note actually documents, and the output surfaces the supporting note text for each code returned.

04

Customization Studio

Per-clinician note preferences applied from the first note and refined from clinician edits over time.

99.92% note approval rating 89% clinician adoption www.deepscribe.ai

Mapped capabilities

4 capabilities

  • Stated preference adherence

    Configured format, section order, and verbosity preferences are honored in generated notes.

  • Learning from edits

    Repeated clinician corrections are reflected in subsequent notes.

  • Preference vs. clinical-content boundary

    Style customization never removes or alters clinically material content.

  • Preference isolation

    One clinician's preferences do not leak into another clinician's notes.

05

EHR Integration and Embedded Access

Bi-directional sync into EHR discrete fields and in-context use inside OncoEMR and iKnowMed.

Access DeepScribe from inside OncoEMR Ⓡ and iKnowMed platforms. www.deepscribe.ai

Mapped capabilities

4 capabilities

  • Discrete-field mapping

    Note content lands in the correct structured EHR fields, not a single free-text blob.

  • Bi-directional data flow

    Inbound schedule and chart data and outbound note data stay consistent.

  • Patient/encounter binding

    Generated content is attached to the correct patient and correct scheduled encounter.

  • Sync failure and recovery

    Behavior when the EHR is unreachable or a write is rejected — work is preserved, not silently lost.

06

Safety, Privacy, and Compliance

The trust envelope around an ambient system operating on PHI in a clinical setting.

the DeepScribe Ambient Operating System is in more than 1,500 healthcare organizations www.deepscribe.ai

Mapped capabilities

4 capabilities

  • PHI handling boundaries

    Patient identifiers are not exposed outside the authorized clinical context.

  • Uncertainty and abstention

    Ambiguous or inaudible content is flagged for clinician review rather than guessed.

  • Clinician-in-the-loop review

    Generated notes and codes are presented as drafts requiring approval before they become the record.

  • Consent and recording state

    Capture only occurs in an explicitly started encounter state.

Coverage is mapped from DeepScribe's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for DeepScribe test?+

The coverage map is generated from DeepScribe's own public product surface (ambient AI clinical documentation for oncology and specialty care): 6 scoring areas — Ambient Clinical Documentation, SmartPrep Pre-Visit Intelligence, and AI Coding, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the DeepScribe evals scored?+

Every case generated for DeepScribe — across Ambient Clinical Documentation and SmartPrep Pre-Visit Intelligence and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the DeepScribe library include?+

The full DeepScribe library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Encounter-to-note fidelity and Oncology-specific structure under Ambient Clinical Documentation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against DeepScribe or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped DeepScribe areas and set them up in a Corsac workspace, where you can run every test case against DeepScribe or your own agent with your own data.