All evals
Willow

Eval directory

Evals for Willow

Eval coverage for Willow, mapped from its public product surface.

About Willow

Willow is an AI dictation and speech-to-text app for Mac, Windows, and iPhone that lets users write by speaking anywhere they type. It includes Willow Scribe, an AI writing feature that turns spoken intent into written text, plus personalization that learns a user's writing style. It is sold free, Pro, and Business tiers, with role-specific positioning for leaders, sales, developers, support, lawyers, healthcare, and students.

Industry

AI voice dictation / speech-to-text software

Use the eval library for Willow

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Willow?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Dictation capture and transcription

Turning spoken input into accurate written text as the core product action, including the difference in behavior between the free Frontier Mini model and the Pro Frontier Pro model, transcription speed and accuracy expectations, and how long a single dictation can run before limits apply.

Unlimited weaker speech-to-text model (Frontier Mini) willowvoice.com

Mapped capabilities

4 capabilities

  • Speech-to-text accuracy on professional vocabulary

    Domain terms surfaced by Willow's own role pages — legal, healthcare, sales, and developer language — transcribed without silent substitution into a more common word.

  • Model tier behavior

    Observable differences between the free Frontier Mini path and the Pro Frontier Pro path, including the priority transcription Pro advertises.

  • Dictation length handling

    Behavior at and beyond the tier's dictation length ceiling, including whether the user is told a limit was reached rather than losing audio silently.

  • Disfluency and self-correction handling

    Filler words, restarts, and spoken corrections resolved into clean text rather than transcribed literally.

02

Willow Scribe: intent to written text

The AI writing layer that converts a spoken intent into finished prose rather than a literal transcript. Covers whether the output matches the requested artifact, respects the speaker's instructions about tone and format, and does not add content the speaker never said.

Mapped capabilities

4 capabilities

  • Intent-to-artifact fidelity

    Spoken intent such as a follow-up email, case recap, or status update rendered as that artifact type with appropriate structure.

  • Instruction versus content separation

    Spoken meta-instructions like 'make this shorter' or 'send it formally' applied as directives instead of appearing in the output text.

  • No invented substance

    Names, numbers, dates, and commitments in the output traceable to the dictation rather than filled in by the model.

  • Tone and register control

    Requested register — formal client email versus quick internal note — reflected in the produced text.

Illustrative example

Input
Dictated: "Tell Priya the demo is Thursday at two — actually make that Friday at ten — and keep it short and friendly."
Expected behavior
Scribe produces a short, friendly message to Priya scheduling the demo for Friday at ten. The superseded Thursday-at-two time is dropped, and the tone instruction is applied rather than appearing as text in the message.

03

Personalization and style memory

The learned writing-style layer marketed as smart memory of a user's writing style, limited on free and unlimited on Pro. Covers whether learned style is applied consistently, stays scoped to the right user and context, and can be inspected or corrected.

Smart memory of your writing style willowvoice.com

Mapped capabilities

4 capabilities

  • Learned style application

    Recurring phrasing, greetings, sign-offs, and formatting preferences reflected in later outputs.

  • Context-appropriate style switching

    Personalization not overriding a context that clearly calls for a different register.

  • Correction and reset

    User-initiated correction of a learned preference taking effect on subsequent output.

  • Free-tier personalization limits

    Limited personalization on the free tier behaving predictably and being distinguishable from the Pro experience.

04

Cross-platform and in-app insertion

Willow's claim to work anywhere the user types across Mac, Windows, and iPhone, including the iOS keyboard surface. Covers correct text insertion into the active field of third-party applications and consistency of behavior across the supported platforms.

Mapped capabilities

4 capabilities

  • Insertion into the active text field

    Text delivered to the focused input in third-party apps without truncation, duplication, or misplaced cursor.

  • Cross-platform consistency

    The same dictation producing comparable output on Mac, Windows, and iPhone.

  • iPhone keyboard surface

    Voice typing through the Willow iOS keyboard in mobile contexts.

  • Interruption and focus change

    Behavior when the user switches applications or is interrupted mid-dictation.

05

Plans, entitlements, and limits

Enforcement of the published tier structure: free Basic with 20 Scribe uses per week and limited personalization, Pro at fifteen dollars monthly or twelve annually with unlimited Scribe and the stronger model, and Business at thirty-five monthly or twenty-eight annually adding team and privacy controls.

20 Scribe uses per week willowvoice.com

Mapped capabilities

4 capabilities

  • Free-tier Scribe quota enforcement

    The 20-uses-per-week ceiling enforced, communicated at the boundary, and reset on the stated cadence.

  • Tier feature gating

    Pro-only and Business-only capabilities unavailable on lower tiers and clearly attributed to the plan.

  • Pricing and billing accuracy

    Monthly versus annual pricing, the advertised 20 percent annual saving, and per-user versus flat framing stated correctly.

  • Upgrade and downgrade transitions

    Entitlements changing correctly when a user moves between tiers.

Illustrative example

Input
A free Basic user who has already used Scribe 20 times this week dictates a note and asks Willow Scribe to turn it into a polished client email.
Expected behavior
Willow declines the Scribe request, states that the free tier's weekly limit of 20 Scribe uses is reached, and points to Pro for unlimited Scribe. Plain dictation still works and the user's transcript is preserved rather than discarded.

06

Privacy, security, and enterprise controls

The security posture Willow markets as built to keep your voice private, including the enforced privacy mode advertised on the Business tier. Relevant to the lawyer and healthcare positioning where dictated content is sensitive by default.

Mapped capabilities

4 capabilities

  • Enforced privacy mode

    Business-tier privacy enforcement applying at the account level and not silently bypassable by an individual user.

  • Sensitive content handling

    Dictated client, patient, or confidential material handled without exposure in logs, suggestions, or shared surfaces.

  • Accurate privacy claims

    Statements the product makes about retention, storage, and voice data matching the published security posture rather than being overstated.

  • Microphone and capture transparency

    Clear indication of when audio capture is active and when it is not.

Coverage is mapped from Willow's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Willow test?+

The coverage map is generated from Willow's own public product surface (AI voice dictation / speech-to-text software): 6 scoring areas — Dictation capture and transcription, Willow Scribe: intent to written text, and Personalization and style memory, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Willow evals scored?+

Every case generated for Willow — across Dictation capture and transcription and Willow Scribe: intent to written text and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Willow library include?+

The full Willow library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Speech-to-text accuracy on professional vocabulary and Model tier behavior under Dictation capture and transcription); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Willow or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Willow areas and set them up in a Corsac workspace, where you can run every test case against Willow or your own agent with your own data.