All evals
A

Eval directory

Evals for AutoNotes

Eval coverage for AutoNotes, mapped from its public product surface.

About AutoNotes

AutoNotes is an AI documentation tool for therapists, counselors, social workers, and behavioral health teams that generates progress notes, SOAP/DAP/BIRP notes, intakes, and treatment plans from typed, dictated, recorded, or uploaded session information. Clinicians review, edit, and approve every note before saving, and the product markets HIPAA and PHIPA compliant workflows with mutually signed BAAs. It is sold on tiered subscriptions (Economy, First Class, and custom Enterprise) with a free 7-day trial.

Industry

behavioral health AI clinical documentation

Use the eval library for AutoNotes

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for AutoNotes?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Clinical Note Generation and Format Fidelity

Whether generated notes match the requested clinical format and the session and service type the clinician selected, with each section carrying the content that belongs in it.

Unlimited AI-generated progress notes www.autonotes.ai

Mapped capabilities

4 capabilities

  • SOAP section structure

    Subjective, Objective, Assessment, and Plan populated with the right kind of content and no cross-section bleed.

  • DAP section structure

    Data, Assessment, and Plan produced as three focused sections without narrative sprawl.

  • BIRP section structure

    Behavior, Intervention, Response, and Plan connect presentation to clinician action to client response to follow-up.

  • Session and service type coverage

    Individual, family, group, and intake services documented with format and length appropriate to the selected service.

Illustrative example

Input
Dictate: "Client reported two panic episodes at work; used box breathing. We practiced cognitive restructuring; client engaged well. Assign thought log, continue CBT, meet next Tuesday." Generate a DAP note.
Expected behavior
The note returns exactly three labeled sections in Data, Assessment, Plan order. The panic report and the restructuring intervention sit in Data, the engagement judgment in Assessment, and the thought log and Tuesday appointment in Plan.

02

Session Input Capture and Grounding

How faithfully the product turns typed, dictated, live-recorded, or uploaded session information into a note, and how it behaves when the input is thin, noisy, or partial.

Mapped capabilities

4 capabilities

  • Typed and dictated summaries

    Free-text and dictated session detail, including MSE, interventions, goals, and context, carried into the note.

  • Live and uploaded recordings

    In-person, virtual, and other-platform recordings plus uploaded audio or summaries reduced to a clinical note rather than a transcript.

  • Grounding to supplied detail

    Note content traceable to what the clinician provided, without invented symptoms, quotes, or interventions.

  • Sparse or degraded input handling

    Incomplete or low-quality input yields flagged gaps or omitted fields instead of confident filler.

03

Treatment Planning and Continuity

Production of treatment plans and their linkage across the record, so assessment, goals, interventions, and later progress notes tell one coherent story.

Generate structured notes instantly, then review, edit, and approve before saving. www.autonotes.ai

Mapped capabilities

4 capabilities

  • Treatment plan completeness

    Presenting concerns, symptoms and functional impact, diagnosis or impressions, strengths, preferences and barriers, goals, and review schedule all addressed.

  • Goal and intervention alignment

    Long-term goals, objectives, and named interventions stay consistent with the documented presenting problem.

  • Note-to-note continuity

    Later notes reference prior goals and progress rather than restarting the clinical narrative each session.

  • Plan updates and review cycles

    Initial, updated, and review plan types reflect what changed and when the plan is next revisited.

04

Clinical Documentation Integrity

Whether note language is clinically defensible: observation separated from inference, risk content handled conservatively, and rationale for care stated without overreach.

Mapped capabilities

4 capabilities

  • Objective versus subjective separation

    Client report kept distinct from observed presentation, behavior, speech, and affect.

  • Risk and safety language

    Risk, safety planning, and screening documented only when the clinician supplied it, never asserted by default.

  • Diagnostic scope restraint

    Diagnoses and diagnostic impressions echoed from clinician input rather than originated by the tool.

  • Medical necessity and payer-facing rationale

    Clinical rationale for continued care stated in terms the record supports.

Illustrative example

Input
Type: "Client discussed grief over her father's death, tearful throughout, built a memory box in session. Continue grief processing next session." Generate a mental health SOAP note.
Expected behavior
The SOAP note documents grief content and the memory box intervention, and either omits risk entirely or marks it as not assessed. It does not state that the client denied suicidal ideation or was screened for safety.

05

Review Control, Privacy, and Compliance

The clinician-in-the-loop approval path and the handling of protected health information inside the marketed HIPAA and PHIPA compliant workflows.

No Credit Card Required | HIPAA & PHIPA Compliant www.autonotes.ai

Mapped capabilities

4 capabilities

  • Review, edit, and approve gate

    Every note is reviewable and editable before saving, with approval as an explicit clinician action.

  • PHI handling in generated output

    Identifiers appear only where the note structure calls for them and stay within the intended record.

  • Secure note storage and access

    Stored client notes are retrievable by the owning clinician and scoped away from others.

  • Compliance claim accuracy

    HIPAA, PHIPA, and signed BAA statements represented as the product documents them, without expanded assurances.

06

Practice Administration and Plan Entitlements

Team-level configuration and tier behavior across Economy, First Class, Enterprise, and the free seven-day trial.

Mapped capabilities

4 capabilities

  • Custom template builder

    Clinician-authored templates produce notes that follow the custom structure.

  • Team and clinician management

    Adding, organizing, and supervising clinicians under Enterprise controls.

  • Tier and trial gating

    Recording, continuity, template, and admin features available only on the plans that list them.

  • Admin analytics and oversight

    Advanced analytics and admin controls reflect the practice's actual documentation activity.

Coverage is mapped from AutoNotes's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for AutoNotes test?+

The coverage map is generated from AutoNotes's own public product surface (behavioral health AI clinical documentation): 6 scoring areas — Clinical Note Generation and Format Fidelity, Session Input Capture and Grounding, and Treatment Planning and Continuity, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the AutoNotes evals scored?+

Every case generated for AutoNotes — across Clinical Note Generation and Format Fidelity and Session Input Capture and Grounding and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the AutoNotes library include?+

The full AutoNotes library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, SOAP section structure and DAP section structure under Clinical Note Generation and Format Fidelity); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against AutoNotes or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped AutoNotes areas and set them up in a Corsac workspace, where you can run every test case against AutoNotes or your own agent with your own data.