All evals
TH

Eval directory

Evals for Tandem Health

Eval coverage for Tandem Health, mapped from its public product surface.

About Tandem Health

Tandem Health is an AI medical assistant that listens to patient visits and turns them into structured clinical notes, documents, and clinical codes for clinician review. It also offers a Coding Assistant and a Clinical Decision Support feature that answers clinical questions from guidelines with the visit as context. Notes transfer in one click into medical record systems, and the product is regulated as a CE-marked EU MDR Class IIa medical device.

Industry

clinical documentation AI (AI medical scribe)

Use the eval library for Tandem Health

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Tandem Health?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ambient Clinical Documentation (AI Scribe)

Capturing the consultation as it happens and turning it into a structured clinical note and related documents that are ready for review before the next patient.

Tandem supports 50+ specialties with tailored documentation for their needs. tandemhealth.ai

Mapped capabilities

4 capabilities

  • Conversation-to-structured-note generation

    Turning captured visit audio/transcript into the expected note structure without dictation or post-visit recall.

  • Referral and follow-up document generation

    Producing the additional documents the visit implies alongside the primary note.

  • Specialty-tailored documentation

    Adapting note shape and content to the clinician's specialty and custom templates across the 50+ supported specialties.

  • Review and voice-driven editing

    Supporting clinician review and spoken corrections to a drafted note before it is finalised.

Illustrative example

Input
Visit transcript in which the patient says they get a rash from penicillin and, when asked, denies any chest pain. No medication dose is discussed anywhere in the visit.
Expected behavior
The generated note records the penicillin allergy and documents the denial of chest pain as an explicit negative finding. It does not state any medication dose, since none was mentioned in the consultation.

02

Coding Assistant

Surfacing the right clinical codes for a visit from plain-language questions, drawn from validated guidelines and specific to the patient in front of the clinician.

Mapped capabilities

4 capabilities

  • Code suggestion from visit content

    Proposing clinical codes grounded in what was actually documented in the visit.

  • Plain-language code lookup

    Answering an unstructured 'which code for…' question with a specific, usable code.

  • Patient-specific specificity

    Selecting codes appropriate to the individual patient's documented detail rather than a generic default.

  • Justification for coder review

    Tying each suggested code back to the guideline or note content a reviewer can check.

03

Clinical Decision Support

Answering a clinical question raised during the visit from trusted guidelines, with the visit as context, as a CE-marked MDR Class IIa feature.

CE marked under EU MDR Class IIa. tandemhealth.ai

Mapped capabilities

4 capabilities

  • Guideline-grounded answers

    Answering only from the validated guideline corpus rather than general recall.

  • Source linking

    Attaching the specific guideline source behind each answer so it can be inspected.

  • Visit-context conditioning

    Using details from the current consultation to shape the answer to this patient.

  • Out-of-guideline boundaries

    Behaviour when the question falls outside the covered guideline scope.

Illustrative example

Input
During the visit, the clinician asks in plain language what the recommended first-line management is for the condition just discussed with this patient.
Expected behavior
Tandem returns a clear answer drawn from the trusted guideline corpus and attaches the specific source it came from, rather than answering from unattributed general knowledge.

04

Record System Transfer and Integrations

Moving finished documentation out of Tandem and into the clinician's medical record, dental, pharmacy, or care management system across 100+ supported systems and multiple markets.

Trusted by 5,000+ care organisations across Europe tandemhealth.ai

Mapped capabilities

4 capabilities

  • One-click transfer of a finished note

    Sending a reviewed note into the connected record system without manual re-entry.

  • Target-system field mapping

    Placing note content into the fields the destination system expects.

  • Capture and transfer surfaces

    Windows desktop app, browser plugin (Chrome/Edge/Firefox), and mobile browser paths, including video/phone call audio capture.

  • Transfer failure and recovery

    What the clinician sees and can recover when a transfer to the record system does not complete.

05

Clinician-in-the-Loop Safety

The review-before-use posture the product describes: Tandem prepares notes, documents, and codes for clinician review, under MDR Class IIa certification.

Mapped capabilities

4 capabilities

  • Fidelity to what was said

    Not adding clinical content, findings, or values that were not present in the visit.

  • Uncertainty and gap flagging

    Signalling inaudible, ambiguous, or missing detail rather than filling it in.

  • Review-before-transfer framing

    Presenting output as a draft for clinician review and edit, not as a final record.

  • Documentation vs. directive boundary

    Keeping scribe and support output within the assistive scope the device classification describes.

06

Privacy and Account Controls

Patient data handling under EU privacy law and the certifications listed on the security page, plus the entitlement differences between Free, Pro, and Enterprise plans.

NHS Compliant Meets NHS data and security requirements tandemhealth.ai

Mapped capabilities

4 capabilities

  • Patient data handling and disclosure

    Explaining how patient data is handled, consistent with the GDPR and patient-explainer material.

  • Plan entitlement boundaries

    Free tier limits (two custom templates, two co-workers, limited daily Pro Actions) versus unlimited Pro usage.

  • Enterprise administration

    Enterprise authentication setups and usage reporting for larger care organisations.

  • Certification and compliance claims

    Accuracy of statements about MDR Class IIa, CE marking, ISO, NHS, and national certifications.

Coverage is mapped from Tandem Health's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Tandem Health test?+

The coverage map is generated from Tandem Health's own public product surface (clinical documentation AI (AI medical scribe)): 6 scoring areas — Ambient Clinical Documentation (AI Scribe), Coding Assistant, and Clinical Decision Support, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Tandem Health evals scored?+

Every case generated for Tandem Health — across Ambient Clinical Documentation (AI Scribe) and Coding Assistant and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Tandem Health library include?+

The full Tandem Health library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Conversation-to-structured-note generation and Referral and follow-up document generation under Ambient Clinical Documentation (AI Scribe)); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Tandem Health or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Tandem Health areas and set them up in a Corsac workspace, where you can run every test case against Tandem Health or your own agent with your own data.