All evals
S

Eval directory

Evals for Scribeberry

Eval coverage for Scribeberry, mapped from its public product surface.

About Scribeberry

Scribeberry is an ambient AI medical scribe that listens to clinical encounters and generates structured notes, letters, and auto-filled forms for clinicians in Canada and the US. It works with any web-based EMR via a Chrome extension, desktop, and mobile apps, and offers 2,000+ templates, 40+ language support, and pre-intake AI agents. It also exposes an API and SDK so EHR vendors, telehealth platforms, and healthcare startups can embed transcription and note generation, with HIPAA/PIPEDA compliance, SOC 2 Type 2 certification, and region-specific data hosting.

Industry

ambient AI medical scribe / clinical documentation AI

Use the eval library for Scribeberry

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Scribeberry?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ambient Note Generation

Turning a live or uploaded clinical encounter into a structured, faithful note in the requested format (SOAP, H&P, consult notes, letters), including pre-chart preparation and supplied clinical context.

Ambient AI clinical notes, letters, and forms from patient conversations. scribeberry.com

Mapped capabilities

4 capabilities

  • Encounter-to-note generation

    Ambient listening or uploaded transcript produces a complete structured note for the stated visit type.

  • Format and specialty conformance

    Output matches the requested note type, specialty, and template structure rather than a generic summary.

  • Grounding and omission control

    Findings, medications, and plans in the note trace back to the encounter; nothing clinically material is dropped or invented.

  • Clinical context and pre-chart inputs

    Clinician-supplied medical context, prior history, and pre-chart review are incorporated into the generated note.

Illustrative example

Input
Transcript: "Patient reports two days of cough, no fever. Continue lisinopril 10 mg daily." Generate a SOAP note using the family-medicine template.
Expected behavior
Cough and the absence of fever appear under Subjective and the lisinopril 10 mg daily instruction under Plan, with all four SOAP headings present. No vitals, exam findings, or diagnoses appear that the transcript did not state.

02

Forms, Templates & Intake Agents

Reusable documentation assets and the automation that fills them: the template library, macros, PDF and referral auto-fill including multi-page forms, and pre-intake AI agents that gather patient information before the visit.

Mapped capabilities

4 capabilities

  • Instant form auto-fill

    New PDFs and referral forms are completed from encounter context without manual prep.

  • Multi-page and specialty form mapping

    Fields across multi-page and specialty-specific forms are mapped to the correct values.

  • Template library and custom templates

    Selection, creation, and consistent application of templates from the 2,000+ library, including drop-in-note behavior.

  • Macros and pre-intake agents

    Macro expansion and pre-intake agent capture feed correctly into the resulting note or form.

03

EMR & Workflow Integration

Getting documentation into the systems clinicians already use — any web-based EMR via the Chrome extension, named one-click workflows, telehealth platforms, and the desktop, mobile, and web clients.

Mapped capabilities

4 capabilities

  • One-click transfer to EMR

    Smart pull and push place note content into the correct EMR fields without tab switching.

  • Cross-EMR coverage

    Named one-click workflows (Accuro, Oscar Pro, Epic, Jane) and generic web-EMR support behave consistently.

  • Telehealth and in-person capture

    Encounters from Zoom, Teams, Google Meet, and in-room visits are captured through the same workflow.

  • Cross-device continuity

    Web, iOS, Android, and Chrome extension surfaces stay in sync for the same encounter.

04

Multilingual Capture & Dictation

Speech handling beyond the default case: multilingual encounters, live translation, natural dictation with medical vocabulary, and uploaded or batch audio for post-visit documentation.

Stream audio and receive medical-grade transcripts in real time. Supports 80+ languages scribeberry.com

Mapped capabilities

4 capabilities

  • Multilingual encounter handling

    Non-English and mixed-language encounters are transcribed and rendered into the requested note language.

  • Live translation

    Translated output preserves clinical meaning, medications, and dosages.

  • Medical-grade dictation

    Dictated speech, including dictation into an already-completed note, is transcribed with correct clinical terminology.

  • Uploaded and async audio

    Recorded files processed after the visit produce the same note quality as live capture.

05

Privacy, Security & Compliance Behavior

How the product describes and enforces its published data commitments: no stored audio, temporary encrypted note storage, no model training on user data, region-specific hosting, and HIPAA/PIPEDA/SOC 2 posture.

HIPAA and PIPEDA compliant. SOC 2 Type 2 certified. Data is encrypted in transit and at rest. scribeberry.com

Mapped capabilities

4 capabilities

  • Data handling claim accuracy

    In-product answers about storage, encryption, and training match the published security policy.

  • Audio retention and deletion

    Zero audio retention and deletion of synced notes behave as stated.

  • Regional data residency

    Canadian and US sessions are described and routed to their respective regional hosting.

  • Compliance representations

    HIPAA, PIPEDA, SOC 2 Type 2, and BAA statements are represented without overclaiming.

Illustrative example

Input
A clinician asks in-product: "Where is my audio stored, and do you train your models on my patients' notes?"
Expected behavior
The reply states that no audio recordings are stored or created, that notes are encrypted and only temporarily retained to sync across devices, that models are never trained on user data, and that Canadian data stays in Canada.

06

Developer API & Plan Entitlements

The embeddable surface for EHR vendors, telehealth platforms, and healthcare startups — realtime and async transcription, note generation, template management via the SDK — plus the plan tiers and usage limits that gate access.

Mapped capabilities

4 capabilities

  • SDK note generation contract

    notes.generate honors transcript, template, and specialty parameters and returns the documented shape.

  • Realtime streaming transcription

    Streamed audio yields incremental clinical transcripts with context-aware terminology.

  • Async batch transcription

    Uploaded recordings are processed for post-visit, dictation, and telehealth workflows.

  • Plan gating and usage limits

    Trial (20 uses/month after 3 days), Pro, and Enterprise entitlements such as API access and custom integrations are enforced as published.

Coverage is mapped from Scribeberry's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Scribeberry test?+

The coverage map is generated from Scribeberry's own public product surface (ambient AI medical scribe / clinical documentation AI): 6 scoring areas — Ambient Note Generation, Forms, Templates & Intake Agents, and EMR & Workflow Integration, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Scribeberry evals scored?+

Every case generated for Scribeberry — across Ambient Note Generation and Forms, Templates & Intake Agents and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Scribeberry library include?+

The full Scribeberry library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Encounter-to-note generation and Format and specialty conformance under Ambient Note Generation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Scribeberry or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Scribeberry areas and set them up in a Corsac workspace, where you can run every test case against Scribeberry or your own agent with your own data.