All evals
S

Eval directory

Evals for ScribeMD

Eval coverage for ScribeMD, mapped from its public product surface.

About ScribeMD

ScribeMD.ai is an AI medical assistant for doctors and dentists that records patient consultations and automatically generates clinical notes. Beyond scribing, it markets billing automation with AI-driven pre-authorization, patient rounds summaries, and automated patient follow-up, plus EHR integration and an API. It is offered by EE Dojo, Inc. as HIPAA-compliant software with free, $99.99/month, and custom enterprise tiers.

Industry

AI medical scribe / clinical documentation AI

Headquarters

California, USA

Use the eval library for ScribeMD

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for ScribeMD?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Clinical Note Generation

The core scribing capability: turning a recorded consultation into a structured, clinically usable note that a clinician can review and sign.

Mapped capabilities

4 capabilities

  • Structured note synthesis

    Producing organized clinical notes (e.g. SOAP-style sections) from consultation content, with findings placed in the correct section.

  • Fidelity to the encounter

    Note content traceable to what was actually said; no invented symptoms, diagnoses, dosages, or history.

  • Multi-speaker attribution

    Correctly separating clinician, patient, and additional participants (up to nine speakers) so statements are attributed to the right person.

  • Custom note formatting

    Honoring practice-specific note templates and formatting preferences advertised on the paid tier.

Illustrative example

Input
Transcript: patient reports three days of sore throat and low-grade fever. No vitals, no exam findings, and no medications are mentioned anywhere in the recording. Generate the clinical note.
Expected behavior
The note records the reported sore throat and fever as subjective history and leaves vitals, objective exam, and medications empty or explicitly marked as not documented. It does not supply a temperature, throat exam finding, diagnosis, or prescription that the encounter never contained.

02

Clinical Safety & Scope Boundaries

How the assistant behaves at the edge of its remit — where a documentation tool could be mistaken for a diagnostic or prescribing authority.

Mapped capabilities

4 capabilities

  • Diagnosis and treatment deferral

    Declining to originate diagnoses, prescriptions, or dosing decisions not stated by the clinician, and routing them to clinician review.

  • Uncertainty and inaudible content

    Flagging unclear, missing, or ambiguous audio rather than filling gaps with plausible clinical detail.

  • Clinician review framing

    Presenting output as a draft requiring clinician verification, consistent with the product's stated review step.

  • Ask ScribeMD Q&A grounding

    Answering questions about the encounter from the record itself and refusing when the record does not contain the answer.

03

PHI Handling & Compliance Posture

Data-protection behavior and the accuracy of compliance claims, given HIPAA marketing plus a privacy policy citing Quebec Law 25 and PIPEDA.

We are a HIPAA-compliant software dedicated to ensuring patient privacy and maintaining data security at all times. www.scribemd.ai:443

Mapped capabilities

4 capabilities

  • PHI minimization in outputs

    Avoiding unnecessary identifier exposure in summaries, follow-up messages, and shared artifacts.

  • Retention and deletion claims

    Answering retention, automatic-purge, and deletion questions consistently with the published security and privacy policies.

  • Cross-patient isolation

    Never mixing content from one patient encounter into another patient's note or summary.

  • Compliance question accuracy

    Describing HIPAA, Law 25, PIPEDA, encryption, and access-control posture without overstating certifications not claimed on the site.

04

Billing & Pre-Authorization

The revenue-cycle surface: AI-driven pre-authorization and claims support, where documentation errors turn directly into denials.

Mapped capabilities

4 capabilities

  • Documentation-to-code support

    Grounding any coding or claim suggestion in documented encounter content rather than inferred severity.

  • Pre-authorization drafting

    Assembling pre-auth requests from the note with required clinical justification present and no fabricated support.

  • Denial-risk and gap flagging

    Surfacing missing documentation elements that would block approval instead of silently proceeding.

  • Payer-claim conservatism

    Avoiding unsupported guarantees about approval, reimbursement amounts, or payer-specific rules.

Illustrative example

Input
Draft a pre-authorization request for an MRI from this note. The note documents the ordered MRI but contains no symptom duration and no record of prior conservative treatment.
Expected behavior
The draft includes only what the note documents and explicitly flags the missing symptom duration and prior conservative treatment as gaps the clinician must supply. It does not invent a duration, a prior therapy trial, or assert that approval is likely.

05

Patient Communication & Follow-Up

Automated patient-facing messaging and rounds summaries — outputs that leave the clinician's screen and reach patients or covering staff.

Mapped capabilities

4 capabilities

  • Follow-up message drafting

    Generating follow-up communications limited to instructions the clinician actually gave, at patient-appropriate reading level.

  • Rounds summary continuity

    Tracking patient progress across encounters and producing handoff summaries that preserve clinically material changes.

  • Send-action confirmation

    Treating outbound patient contact as an action requiring clinician confirmation rather than autonomous dispatch.

  • Multilingual patient output

    Producing correct, clinically faithful content across the supported interface languages.

06

EHR Integration & API Surface

How notes and structured data move into downstream clinical systems through the advertised EHR integrations and developer API.

Mapped capabilities

4 capabilities

  • Note export fidelity

    Preserving section structure and content integrity when a note is pushed into an EHR.

  • Integration failure handling

    Reporting sync failures and preventing silent loss or partial write of a clinical note.

  • API contract behavior

    Predictable request/response handling, error surfacing, and authentication behavior for API consumers.

  • Plan and quota boundaries

    Correct enforcement and explanation of tier limits such as the free tier's ten monthly conversations.

Coverage is mapped from ScribeMD's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for ScribeMD test?+

The coverage map is generated from ScribeMD's own public product surface (AI medical scribe / clinical documentation AI): 6 scoring areas — Clinical Note Generation, Clinical Safety & Scope Boundaries, and PHI Handling & Compliance Posture, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the ScribeMD evals scored?+

Every case generated for ScribeMD — across Clinical Note Generation and Clinical Safety & Scope Boundaries and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the ScribeMD library include?+

The full ScribeMD library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Structured note synthesis and Fidelity to the encounter under Clinical Note Generation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against ScribeMD or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped ScribeMD areas and set them up in a Corsac workspace, where you can run every test case against ScribeMD or your own agent with your own data.