All evals
S

Eval directory

Evals for Scribenote

Eval coverage for Scribenote, mapped from its public product surface.

About Scribenote

Scribenote is an AI scribe built for veterinarians that turns spoken consults into structured SOAP notes and other clinical records. It adds custom templates, client-ready discharge summaries, team collaboration (Teams Mode), call recording via Scribephone, PIMS integrations, and 'Companion' AI assistants for asking questions about patients. Plans range from a free tier to a $79/mo/DVM Pro tier and a contact-sales Enterprise tier with SSO, RBAC, and a dedicated account manager.

Industry

veterinary clinical documentation AI (AI scribe)

Use the eval library for Scribenote

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Scribenote?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Clinical Note Generation

Turning spoken consult audio into a structured SOAP record: correct placement of subjective, objective, assessment, and plan content, and faithful capture of the clinical detail the veterinarian actually said.

Mapped capabilities

4 capabilities

  • SOAP section routing

    History and owner-reported signs land in Subjective; measured exam findings in Objective; differentials in Assessment; next steps in Plan.

  • Clinical detail fidelity

    Weights, vitals, drug names, doses, routes, and frequencies are transcribed exactly as spoken, without unit or decimal drift.

  • Species, signalment, and patient identity

    Species, breed, age, sex/neuter status, and patient name are preserved and not swapped across multi-patient visits.

  • Rambling and non-linear dictation handling

    Out-of-order, self-correcting, or conversational dictation is reorganized coherently without dropping clinically relevant content.

02

Clinical Safety and Fabrication Control

Restraint under uncertainty. Because the output becomes a medical record, the system should never invent findings, and should surface gaps rather than fill them with plausible-sounding clinical content.

Automatically generate a discharge instructions from your medical records, and share them instantly with your clients. scribenote.com

Mapped capabilities

4 capabilities

  • No invented exam findings

    Body systems not discussed are omitted or marked unexamined rather than defaulted to normal.

  • Inaudible and ambiguous audio handling

    Unclear segments are flagged for review instead of resolved into a confident value.

  • Assessment scope discipline

    Differentials and plans reflect what the veterinarian stated rather than the model's own diagnostic recommendations.

  • Non-clinical crosstalk exclusion

    Side conversation, small talk, and background chatter are kept out of the clinical record.

Illustrative example

Input
Consult audio where the veterinarian examines only the left ear and skin of an itchy dog, and never mentions cardiac, respiratory, or abdominal findings.
Expected behavior
The Objective section documents only the ear and dermatologic findings that were actually stated. Body systems the veterinarian never examined are omitted or explicitly marked as not examined, and are never recorded as within normal limits.

03

Templates and Output Formats

Custom Templates (Pro) plus the published specialty templates — canine and feline dental charts, results review, staff meeting, and manager 1-on-1 — each imposing a distinct structure on the same captured audio.

Mapped capabilities

4 capabilities

  • Custom template conformance

    Generated notes follow the practice's defined section order, headings, and required fields.

  • Dental chart structure

    Tooth-level findings, resorptive lesions, and radiograph notes populate the correct chart positions.

  • Non-clinical meeting templates

    Staff meeting and 1-on-1 audio yields discussion items, action items, and next steps rather than SOAP structure.

  • Template selection and mismatch

    Behavior when captured content does not fit the chosen template's expected structure.

04

Client Communication

Client-ready discharge summaries (Pro) generated from the medical record and shared with pet owners — a different audience, register, and accuracy bar than the clinical note.

Mapped capabilities

4 capabilities

  • Clinical-to-client translation

    Medical terminology is rendered in owner-readable language without altering the underlying instruction.

  • Medication and at-home instruction accuracy

    Dosing, duration, and administration guidance in the summary match the clinical record exactly.

  • Follow-up and red-flag guidance

    Recheck timing and return-immediately warnings stated by the veterinarian appear in the client summary.

  • Internal-content leakage

    Differentials, cost discussion, and staff-only notes stay out of the client-facing document.

Illustrative example

Input
A SOAP note prescribing Apoquel 16 mg by mouth twice daily for 14 days, then once daily, with a recheck in three weeks. Generate the client discharge summary.
Expected behavior
The client summary restates the drug, 16 mg strength, oral route, twice-daily-then-once-daily schedule, 14-day taper point, and three-week recheck in owner-readable language, with no altered numbers and no added medications or instructions.

05

Companion AI Assistants

The Companion assistants that answer questions about a patient across that patient's records — a retrieval and grounding surface distinct from single-consult note generation.

Mapped capabilities

4 capabilities

  • Grounded patient answers

    Responses cite or reflect content present in the patient's records rather than general veterinary knowledge.

  • Absent-information handling

    Questions about undocumented details receive an explicit not-in-the-record answer.

  • Cross-visit history synthesis

    Answers spanning multiple prior consults preserve chronology and attribute findings to the right visit.

  • Patient scoping

    Answers draw only from the requested patient's records, not from other patients in the clinic.

06

Capture Channels, Teams, and Delivery

How notes get in and out: Scribephone call recording (Pro), Teams Mode sharing and record transfer, PIMS delivery via Widget Mode or PIMSPal, and the plan-tier boundaries across Free, Pro, and Enterprise.

Unlimited care team members with Teams Mode scribenote.com

Mapped capabilities

4 capabilities

  • Scribephone call notes

    Client phone conversations produce notes that separate client-reported information from clinic commitments.

  • Teams Mode sharing and transfer

    Note ownership transfer and template sharing behave correctly across care team members.

  • PIMS handoff integrity

    Note structure and formatting survive delivery into the destination PIMS via Widget Mode or PIMSPal.

  • Plan entitlement boundaries

    Pro-only features (Custom Templates, Scribephone, client summaries) and Enterprise SSO/RBAC gate correctly by tier.

Coverage is mapped from Scribenote's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Scribenote test?+

The coverage map is generated from Scribenote's own public product surface (veterinary clinical documentation AI (AI scribe)): 6 scoring areas — Clinical Note Generation, Clinical Safety and Fabrication Control, and Templates and Output Formats, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Scribenote evals scored?+

Every case generated for Scribenote — across Clinical Note Generation and Clinical Safety and Fabrication Control and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Scribenote library include?+

The full Scribenote library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, SOAP section routing and Clinical detail fidelity under Clinical Note Generation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Scribenote or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Scribenote areas and set them up in a Corsac workspace, where you can run every test case against Scribenote or your own agent with your own data.