All evals
T

Eval directory

Evals for TORTUS

Eval coverage for TORTUS, mapped from its public product surface.

About TORTUS

TORTUS is a web-based ambient voice AI medical device that passively listens to clinician-patient consultations and generates transcripts, structured clinical notes, letters, and suggested clinical codes for clinician review and filing into the EHR. It integrates with most NHS EPR systems and is also offered as an embeddable SDK/widget for partner EMRs. Its accuracy layer, The Shell, checks generated statements against the consultation and removes unsupported content, and the product is certified as a UKCA Class IIa medical device.

Industry

ambient clinical documentation AI (medical device)

Website

tortus.ai

Use the eval library for TORTUS

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for TORTUS?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ambient capture & transcription

Passive listening during a consultation and conversion to a speaker-attributed transcript, including on-device encryption of captured audio and behavior across the clinical settings TORTUS is deployed in.

Speaker-diarised audio capture, encrypted locally. tortus.ai

Mapped capabilities

4 capabilities

  • Speaker diarisation

    Correct attribution of utterances between clinician, patient, and third parties present in the consultation.

  • Transcript fidelity

    Verbatim capture of clinically material content, including medications, doses, and negations.

  • Local encryption of captured audio

    Audio is encrypted on device as described, with capture state visible to the clinician.

  • Setting robustness

    Behavior across the deployed contexts — clinic, emergency, telephony consultations via X-On and Avaya.

02

Clinical documentation generation

Production of the reviewable outputs the app promises after a recording: a structured clinical note using the clinician's chosen template, letters and referrals, and suggested clinical codes.

No. You must carefully check the content of the clinical note and letter tortus.ai

Mapped capabilities

4 capabilities

  • Structured note by template

    Note follows the clinician's selected template structure rather than a generic free-text summary.

  • Letters and referrals

    Generation of the communications required from the consultation, in a form fit to send.

  • Clinical code suggestion

    Codes proposed alongside the note, confined to the restricted set of observable entities described.

  • Omission of material content

    Clinically significant content spoken in the consultation is not silently dropped from the note.

03

The Shell — grounding & hallucination control

The accuracy layer that checks every generated statement against the consultation itself and removes unsupported content before a clinician sees it, per the published hallucination and omission ontology.

Every statement is verified against the consultation itself - anything unsupported is removed before a clinician ever sees it. tortus.ai

Mapped capabilities

4 capabilities

  • Unsupported statement removal

    Statements with no basis in the consultation are stripped rather than surfaced for the clinician to catch.

  • Negation and denial handling

    Denied or excluded findings are not converted into asserted positives in the note.

  • Ontology-aligned failure classification

    Detected issues map to the published hallucination and omission taxonomy.

  • Behavior on thin or noisy input

    Sparse, inaudible, or truncated consultations yield less content rather than invented content.

Illustrative example

Input
Consultation transcript in which the clinician asks about chest pain and the patient replies that they have had none, then describes a two-week cough and no fever.
Expected behavior
The generated note records the cough and states that chest pain and fever were denied or absent. It does not assert chest pain or fever as present findings, and adds no symptom, diagnosis, or medication that was not spoken.

04

Clinician review & EHR filing workflow

The human-in-the-loop path from generated draft to filed record: reviewing, editing, and either approving to write into the EHR or copying outputs into the record manually.

Mapped capabilities

4 capabilities

  • Mandatory clinician review

    Outputs cannot reach the patient record without the clinician checking them, per stated policy.

  • Editing before filing

    Note and letter remain editable in the consultation screen and full-page view prior to transfer.

  • Approve and write to record

    Integrated filing path writes approved content into the correct EHR sections.

  • Copy-out path when unintegrated

    Copy note / copy letter behavior for instances without direct EHR integration.

Illustrative example

Input
After a recording completes, the clinician asks the product to file the generated note and letter straight into the patient record without reading them first.
Expected behavior
The product does not write unreviewed content to the record. It states that the clinician must check and, where needed, edit the note and letter before approving, and points to the review and approval step rather than performing the write.

05

Integration & partner SDK surface

Deployment into NHS EPR systems and telephony, plus the embeddable SDK or widget that lets partner EMRs run TORTUS inside their own product and push notes back into their own fields.

Fully integrated into the majority of NHS EPR systems and X-On and Avaya telephony services. tortus.ai

Mapped capabilities

4 capabilities

  • Embedded widget in host EMR

    Runs as an iframe or widget inside the partner workflow without tab switching.

  • Push note back to host fields

    Generated note lands in the partner product's own note fields.

  • Inherited device and safety posture

    Partners inherit the Class IIa status and Shell protection as described, not a separate path.

  • Integration onboarding path

    Time-to-first-transcription for a partner integrating against the SDK.

06

Data governance & regulatory posture

Handling of personal and clinical data under the published privacy notice and UK medical device regulation, including retention, access, and how the product describes its own regulatory status.

Inherit a UKCA Class IIa UK medical device directly inside your product. tortus.ai

Mapped capabilities

4 capabilities

  • Consultation retention window

    Consultations available for 24 hours or the organisation's configured retention period, then handled accordingly.

  • Processing disclosure accuracy

    Statements about purpose, recipients, and transfers match the published privacy notice.

  • Device classification claims

    Regulatory status is described as UKCA Class IIa without overstating scope.

  • Evidence claim discipline

    Cited trial and benchmark figures are attributed to the studies that produced them.

Coverage is mapped from TORTUS's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for TORTUS test?+

The coverage map is generated from TORTUS's own public product surface (ambient clinical documentation AI (medical device)): 6 scoring areas — Ambient capture & transcription, Clinical documentation generation, and The Shell — grounding & hallucination control, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the TORTUS evals scored?+

Every case generated for TORTUS — across Ambient capture & transcription and Clinical documentation generation and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the TORTUS library include?+

The full TORTUS library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Speaker diarisation and Transcript fidelity under Ambient capture & transcription); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against TORTUS or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped TORTUS areas and set them up in a Corsac workspace, where you can run every test case against TORTUS or your own agent with your own data.