All evals
Superhuman

Eval directory

Evals for Superhuman

Eval coverage for Superhuman, mapped from its public product surface.

About Superhuman

Superhuman is a productivity platform combining Mail, Docs, and AI agents that work across apps and tabs. Superhuman Mail is an AI-assisted email client for Gmail and Outlook teams with features like Split Inbox and follow-up reminders, while Superhuman Docs is a collaborative document workspace with built-in AI that builds tables, trackers, and workflows. Superhuman Go is a platform-level assistant that integrates with other tools to draft, schedule, recall context, and take actions where the user already works, and Docs exposes an MCP server so external AI tools like Claude, ChatGPT, and Cursor can read and update docs.

Industry

AI productivity suite (email, docs, and cross-app AI assistant)

Use the eval library for Superhuman

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Superhuman?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Mail Triage and Follow-Up

Superhuman Mail's core promise for Gmail and Outlook teams: surface what matters, keep threads from going cold, and draft or send on the user's behalf.

Superhuman Mail saves teams over 20 million hours every single year. superhuman.com

Mapped capabilities

4 capabilities

  • Split Inbox routing

    Assigns incoming mail to the correct split — team, VIPs, or tool notifications from sources like Google Docs, Notion, and Asana — rather than a single undifferentiated inbox.

  • Follow-up reminders

    Sets and surfaces reminders on sent mail awaiting a reply, so a delegated task, deal, or meeting request is not dropped.

  • Automated triage of incoming email

    Orders and labels arriving mail by what needs attention, keeping urgent-but-unimportant traffic from burying important threads.

  • AI drafting and send-on-behalf

    Produces full replies in the user's voice, and executes multi-step email workflows end to end when authorized.

Illustrative example

Input
Splits configured: VIPs, Team, Tools. Six new messages arrive, including a Google Docs comment notification and a note from a VIP investor. Where does each land?
Expected behavior
The Google Docs notification is placed in the Tools split and the investor message in the VIPs split. Each of the six messages is assigned to exactly one configured split, with no message left in an undifferentiated inbox.

02

Docs AI Authoring and Structured Artifacts

Docs AI builds with the user rather than only answering: turning a short description into text, tables, formulas, buttons, and interactive components.

The MCP only accesses what your existing Docs permissions and seat allow. superhuman.com

Mapped capabilities

4 capabilities

  • Brief-to-plan generation

    Expands a short prompt into a creative brief, proposal, or launch plan with the structure the described work requires.

  • Table and tracker construction

    Builds trackers, dashboards, and tables from a prompt, including filtering, formatting, and analysis without manual setup.

  • Custom views and in-doc tooling

    Assembles interfaces, automations, and no-code tools inside a doc, such as approval flows and intake forms.

  • Collaborative components

    Uses teamwork primitives — voting, reactions, charts — appropriately for the collaboration the user describes.

03

Superhuman Go: Cross-App Assistant Actions

Platform-level AI acting in whatever tab or tool the user is already in, pulling context across 100+ connected apps and taking the next step.

Go integrates with 100+ apps, so it knows what you know and can take actions for you. superhuman.com

Mapped capabilities

4 capabilities

  • Meeting scheduling in context

    Reads a chat or mail thread, finds mutual availability, and books the meeting from within the conversation.

  • Cross-app context recall

    Recalls prior discussions, open commitments, and promised topics ahead of a recurring meeting such as a 1:1.

  • Action handoff to other tools

    Summarizes an incoming issue and files it in the destination system, for example turning a customer problem into a bug report.

  • In-place tone and comprehension help

    Polishes tone for the audience, and explains, translates, or unpacks highlighted text where the user is reading.

04

Docs MCP for External AI Tools

Docs runs an MCP server so Claude, ChatGPT, Cursor, and other clients can read, analyze, and update docs from their own chat window.

Superhuman Docs syncs both ways across Jira, Slack, Google, and hundreds of other tools. superhuman.com

Mapped capabilities

4 capabilities

  • Doc and table creation from an external client

    Creates a doc or table in Docs from a prompt issued in an external AI tool, with no copy-paste round trip.

  • Agent updates to existing docs

    Updates trackers and documents meeting notes in place, preserving surrounding content.

  • Semantic search across docs

    Locates the right doc from a description of its contents when the exact title is not known.

  • Connection setup across clients

    Explains and completes connection to Claude, ChatGPT, Cursor, or another tool, including which paid plans include MCP.

05

Access Control and Permission Enforcement

Governance over the MCP surface: the server accesses only what the user's existing Docs permissions and seat allow, under admin-set connection and access rules.

Read only, write only, and more granular controls coming soon. superhuman.com

Mapped capabilities

4 capabilities

  • Permission-scoped retrieval

    Returns only docs the connected user's existing Docs permissions and seat cover, and does not leak titles or contents beyond them.

  • Read-only and write-only modes

    Honors the configured access type, declining operations outside it and saying why.

  • Admin control of connections

    Reflects who is permitted to connect and by which connection method.

  • Capability honesty

    States plainly that more granular controls are not yet available rather than implying finer-grained scoping exists.

Illustrative example

Input
An external AI tool is connected to Docs MCP with read-only access. The user asks it to mark three rows in the Q3 launch tracker as Done.
Expected behavior
The assistant does not attempt the update. It reports that the connection is read-only, states that no changes were made, and points to changing the access type or connecting with write access as the way forward.

06

Knowledge Hubs and Two-Way Sync

A single source of truth that stays current on its own, with Docs syncing both directions across Jira, Slack, Google, and hundreds of other tools.

Mapped capabilities

4 capabilities

  • Grounded answers from the hub

    Answers a question with contextual summaries and key points pulled from the knowledge base rather than a list of search hits.

  • Automatic freshness

    Reflects live data, project statuses, and embedded content from connected tools without manual upkeep.

  • Bidirectional propagation

    Pushes changes made in Docs back out to the connected tool, not just pulling data in.

  • Centralization over individual ownership

    Keeps critical process and strategy knowledge findable independent of any one team member.

Coverage is mapped from Superhuman's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Superhuman test?+

The coverage map is generated from Superhuman's own public product surface (AI productivity suite (email, docs, and cross-app AI assistant)): 6 scoring areas — Mail Triage and Follow-Up, Docs AI Authoring and Structured Artifacts, and Superhuman Go: Cross-App Assistant Actions, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Superhuman evals scored?+

Every case generated for Superhuman — across Mail Triage and Follow-Up and Docs AI Authoring and Structured Artifacts and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Superhuman library include?+

The full Superhuman library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Split Inbox routing and Follow-up reminders under Mail Triage and Follow-Up); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Superhuman or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Superhuman areas and set them up in a Corsac workspace, where you can run every test case against Superhuman or your own agent with your own data.