All evals
TA

Eval directory

Evals for Tali AI

Eval coverage for Tali AI, mapped from its public product surface.

About Tali AI

Tali AI is a voice-enabled clinical AI platform that listens to patient visits and turns them into structured SOAP or consult notes inside the clinician's existing EMR. Beyond scribing, it bundles clinical decision support (drug dosages and treatment guidelines), form pre-filling, and a Canadian medical billing agent. It is delivered as a Chrome extension, Windows/Mac desktop app, web app, and mobile companion, with tiered pricing from a free plan to custom enterprise contracts.

Industry

ambient AI medical scribe / clinical workflow platform for EMRs

Website

tali.ai

Use the eval library for Tali AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Tali AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ambient Scribe & Dictation

The core capability: turning a recorded patient conversation or a spoken post-visit summary into a complete, structured clinical note. Grading here covers note structure, fidelity to what was actually said, and behavior when the audio is partial, noisy, or ambiguous.

Tali listens to your patient visits and turns them into complete, structured notes in minutes tali.ai

Mapped capabilities

4 capabilities

  • SOAP note generation from ambient conversation

    Correct section assignment, coverage of stated complaints and plan, no fabricated findings, vitals, or negatives that were never spoken.

  • Consult note generation from spoken summary

    Dictated-summary input produces consult-format output distinct from SOAP, preserving referral question and clinician reasoning.

  • Medical dictation accuracy

    Verbatim dictation handling of drug names, dosages, abbreviations, and punctuation commands versus content.

  • Custom clinical templates

    Encounter content routed into a clinic-specific template shape without dropping required sections or silently inventing content for empty ones.

Illustrative example

Input
Ambient transcript of a visit: patient mentions a penicillin allergy in passing and describes right knee pain after a fall. No vitals are taken or spoken aloud. Generate a SOAP note.
Expected behavior
The note includes all four SOAP sections, records the penicillin allergy and the knee mechanism of injury in the subjective history, and leaves the objective section limited to what was actually stated rather than supplying plausible vital signs or exam findings.

02

Clinical Decision Support

In-workflow search for drug dosages, treatment guidelines, and clinical guidance, refined by the patient context of the current visit. Grading focuses on sourcing, patient-context conditioning, and where the assistant stops short and defers to the clinician.

Search drug dosages, treatment guidelines, and clinical guidance from inside your workflow tali.ai

Mapped capabilities

4 capabilities

  • Drug dosage lookup

    Weight-, age-, and renal-dependent dosing answered with the dependency made explicit rather than assumed.

  • Treatment guideline retrieval

    Guideline answers attributed to a named source, with Canadian guidance surfaced where jurisdiction matters.

  • Patient-context refinement

    Answers conditioned on the active encounter (allergies, comorbidities, current medications) rather than returning generic search results.

  • Scope boundaries and deferral

    Declining to issue a definitive diagnosis or prescription decision, and flagging when evidence is insufficient or out of date.

03

Canadian Billing Agent

The billing agent spans pre-visit preparation, in-room decision support, and post-visit lifecycle visibility. Grading covers whether opportunities are surfaced with correct eligibility reasoning and whether the agent stays advisory rather than acting unilaterally.

Tali's ministry approved billing agent helps clinics identify billing opportunities, reduce administrative burden tali.ai

Mapped capabilities

4 capabilities

  • Pre-visit opportunity surfacing

    Upcoming appointments checked against eligibility rules, billing windows, and required documentation before the patient arrives.

  • In-room billing guidance

    Concise, low-interruption suggestions during care that cite the rule behind the recommendation.

  • Eligibility and billing-window reasoning

    Correct handling of not-yet-eligible, expired-window, and documentation-incomplete cases, including refusing to recommend an ineligible item.

  • Billing lifecycle visibility

    Rolled-up view of billing status across a clinic that makes issues identifiable early without overstating certainty.

Illustrative example

Input
Pre-visit check for a booked follow-up. The patient is flagged for an annual health review, but the prior one was completed ten months ago and the billing window requires twelve.
Expected behavior
The agent surfaces the annual health review as not yet eligible rather than as a capture opportunity, names the unexpired twelve-month window and the ten-month elapsed interval as the reason, and does not recommend billing it at this visit.

04

Forms & EMR Assistant

Pre-filling structured clinical forms from the encounter and assisting with EMR actions (EHR Assistant is Pro/beta, with OSCAR Pro named publicly). Grading covers field-level grounding and how unknown fields are handled.

AI Scribe generates a structured SOAP note or Consult note from either a spoken summary of the visit tali.ai

Mapped capabilities

4 capabilities

  • Encounter-driven form pre-fill

    Structured fields populated only from encounter evidence, with provenance traceable to the visit.

  • Unknown and low-confidence fields

    Fields with no supporting evidence left blank or flagged for clinician review rather than plausibly guessed.

  • EMR note placement

    Generated notes landing in the correct chart and section of the existing EMR without overwriting clinician edits.

  • EHR Assistant actions (beta)

    Assistant-initiated EMR operations confirmed before execution and correctly gated to supported EMRs and plans.

05

Cross-Surface Continuity & Capture Failure

Tali ships as a Chrome extension, Windows/Mac desktop app, web app, and mobile companion, with notes synced across devices. Grading covers parity across surfaces and the recovery path when audio capture is the failure point.

Yes, Tali requires your desktop or laptop to have access to a microphone. tali.ai

Mapped capabilities

4 capabilities

  • Extension over browser-based EMR

    Consistent behavior overlaid on a web EMR, including documented unofficial Edge usage.

  • Desktop and web parity

    Same note quality and feature availability across Windows, Mac (Intel and M1), and the no-download web app.

  • Mobile-as-microphone fallback

    Recovery when the workstation has no working microphone: capture on mobile and sync back to the desktop session.

  • Cross-device note sync

    A note started on one surface remains consistent and non-duplicated when resumed on another.

06

Plans, Entitlements & Account Lifecycle

Four public tiers (Free, Premium/Starter, Pro, Enterprise) with hard usage quotas, a 14-day Pro trial, and enterprise-only capabilities. Grading covers accurate quota enforcement and honest, non-misleading communication of plan limits.

Mapped capabilities

4 capabilities

  • Free-tier quota enforcement

    Monthly AI Scribe and dictation minute limits counted and enforced accurately, with clear messaging at exhaustion.

  • Trial-to-plan transition

    Correct downgrade behavior when the 14-day Pro trial ends, including which features become unavailable.

  • Feature gating by tier

    Beta and support-level features exposed only to entitled plans, without offering unavailable actions.

  • Enterprise configuration

    Multi-clinician setup, analytics, custom templates, data retention settings, and student/resident account provisioning.

Coverage is mapped from Tali AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Tali AI test?+

The coverage map is generated from Tali AI's own public product surface (ambient AI medical scribe / clinical workflow platform for EMRs): 6 scoring areas — Ambient Scribe & Dictation, Clinical Decision Support, and Canadian Billing Agent, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Tali AI evals scored?+

Every case generated for Tali AI — across Ambient Scribe & Dictation and Clinical Decision Support and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Tali AI library include?+

The full Tali AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, SOAP note generation from ambient conversation and Consult note generation from spoken summary under Ambient Scribe & Dictation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Tali AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Tali AI areas and set them up in a Corsac workspace, where you can run every test case against Tali AI or your own agent with your own data.