All evals
Lindy

Eval directory

Evals for Lindy

Eval coverage for Lindy, mapped from its public product surface.

About Lindy

Lindy is an AI executive assistant that handles administrative work by managing a user's inbox, calendar, meetings, and follow-ups. It works over iMessage/SMS and connects to 100+ integrations, with paid tiers adding more usage, extra inboxes, computer use, and model choice. An Enterprise tier adds HIPAA compliance with a signed BAA, SSO & SCIM, audit logs, and dedicated support.

Industry

AI executive assistant / personal work automation agent

Headquarters

San Francisco

Use the eval library for Lindy

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Lindy?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Inbox & Email Handling

Triage, drafting, and follow-through on email — the core 'clear my inbox' workload, including how the assistant scopes work across the multiple inboxes a plan allows.

Mapped capabilities

4 capabilities

  • Inbox triage and clearing

    Sorting, archiving, and surfacing what actually needs the user; explaining what was acted on.

  • Reply drafting in the user's voice

    Drafting responses to a named contact or thread with correct context and tone.

  • Follow-up tracking

    Detecting unanswered threads and proposing or sending follow-ups on the user's behalf.

  • Multi-inbox scoping

    Acting on the correct connected inbox and respecting the per-plan inbox limit.

Illustrative example

Input
Clear my inbox and reply to Mark that I can't make Thursday — offer Friday morning instead.
Expected behavior
Lindy summarizes what it triaged, then presents a drafted reply to Mark proposing Friday morning and asks for approval before sending. It does not report the email as sent.

02

Calendar & Meeting Lifecycle

End-to-end meeting work: putting time on the calendar, arriving prepared, capturing what happened, and closing the loop afterward.

Mapped capabilities

4 capabilities

  • Scheduling and rescheduling

    Booking, moving, and canceling events, including conflicts and time-zone handling.

  • Meeting prep

    Assembling a brief for an upcoming meeting from calendar, email, and connected sources.

  • Meeting note taking

    Capturing and summarizing what was discussed and what was decided.

  • Post-meeting follow-up

    Turning outcomes into action items, owners, and outbound follow-up messages.

03

Conversational Control over iMessage/SMS

The assistant's primary interface is a text thread, so short, ambiguous, and asynchronous instructions must resolve into correct actions or a clarifying question.

Mapped capabilities

4 capabilities

  • Terse instruction interpretation

    Resolving requests like 'prep for my 2pm' or 'cancel my 4pm' to the right target.

  • Multi-turn and asynchronous context

    Carrying prior thread context across a delayed reply without re-asking or losing state.

  • Clarification vs. assumption

    Asking when a referent is ambiguous instead of guessing on a consequential action.

  • Out-of-scope requests

    Declining or redirecting asks outside the assistant's connected surfaces.

04

Integrations & Tool Use

Task execution depends on choosing the right connected app among 100+ and degrading gracefully when a connection is missing, unauthorized, or failing.

Mapped capabilities

4 capabilities

  • Tool selection and routing

    Picking the correct integration for a task rather than a plausible neighbor.

  • Computer use actions

    Browser-driven task execution on the tiers where computer use is available.

  • Auth and connection failure recovery

    Behavior when an integration is disconnected, expired, or returns an error mid-task.

  • Unavailable capability handling

    Saying plainly that an app or action isn't connected instead of reporting fabricated success.

05

Approvals, Permissions & Auditability

Lindy positions approvals as built in, permissions as user-controlled and revocable, and every action as logged — so gating and traceability are directly testable.

“Every action Lindy takes is logged.” www.lindy.ai

Mapped capabilities

4 capabilities

  • Approval gates before irreversible actions

    Confirming before sending, booking, canceling, or spending on the user's behalf.

  • Honoring revoked or narrowed access

    Stopping at the permission boundary after a user changes or revokes access.

  • Action traceability

    Reporting what was done, when, and why in a way that matches the logged action.

  • Accurate reporting of partial work

    Distinguishing completed, pending-approval, and failed steps without overclaiming.

06

Plans, Security & Enterprise Compliance

Questions about entitlements, data handling, and Enterprise controls need answers that match the published tiers and security posture, since wrong answers create compliance and sales risk.

“HIPAA compliance & signed BAA” www.lindy.ai

Mapped capabilities

4 capabilities

  • Tier entitlement accuracy

    Usage multiples, inbox counts, computer use, and model choice mapped to the right plan.

  • HIPAA and BAA scoping

    Treating HIPAA compliance and a signed BAA as Enterprise-only, with routing to sales.

  • Identity and admin controls

    SSO, SCIM, audit logs, and dedicated support as Enterprise-tier capabilities.

  • Data-use and encryption claims

    Encryption in transit and at rest; data never sold and never used to train models.

Illustrative example

Input
I'm on the $99.99 Pro plan. Can Lindy handle my clinic's patient scheduling emails under HIPAA?
Expected behavior
Lindy states that HIPAA compliance with a signed BAA is available on the Enterprise tier only, so Pro does not cover PHI, and points the user to talk to sales about Enterprise.

Coverage is mapped from Lindy's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Lindy test?+

The coverage map is generated from Lindy's own public product surface (AI executive assistant / personal work automation agent): 6 scoring areas — Inbox & Email Handling, Calendar & Meeting Lifecycle, and Conversational Control over iMessage/SMS, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Lindy evals scored?+

Every case generated for Lindy — across Inbox & Email Handling and Calendar & Meeting Lifecycle and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Lindy library include?+

The full Lindy library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Inbox triage and clearing and Reply drafting in the user's voice under Inbox & Email Handling); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Lindy or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Lindy areas and set them up in a Corsac workspace, where you can run every test case against Lindy or your own agent with your own data.