All evals
Manus

Eval directory

Evals for Manus

Eval coverage for Manus, mapped from its public product surface.

About Manus

Manus is a general AI agent that plans and executes multi-step tasks in a sandboxed virtual computer, delivering finished work products such as slides, websites, designs, research, and deployed apps. It offers web, mobile, and desktop clients, an API, and connectors to tools like Google Calendar, GitHub, Notion, and Slack. A Team plan adds shared workspaces, pooled credits, an admin dashboard, and SSO; the pages state Manus is now part of Meta.

Industry

autonomous general AI agent

Website

manus.im

Use the eval library for Manus

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Manus?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Autonomous task planning and execution

The core agent loop: turning an open-ended request into a plan, executing it across many steps in the sandbox VM, and holding context over a long task. Covers the documented Plan Mode and Branch behaviors and the agent's ability to work without step-by-step supervision.

“Manus AI is an autonomous general AI agent designed to complete tasks and deliver results.” manus.im

Mapped capabilities

4 capabilities

  • Plan formation and Plan Mode review

    Decomposing an ambiguous request into an inspectable plan the user can approve or amend before execution begins.

  • Long-horizon context retention

    Carrying earlier decisions, files, and constraints across many steps of a single task without re-asking or contradicting itself.

  • Parallel and branched exploration

    Running independent subtasks or Branch directions concurrently and reconciling their outputs into one coherent result.

  • Scheduled and recurring tasks

    Interpreting a recurring instruction and producing consistent, correctly scoped output on each scheduled run.

02

Work product generation and fidelity

Manus is positioned on delivering finished artifacts rather than chat answers. This area covers whether the returned slides, sites, designs, images, and audio actually match the brief in format, structure, and factual grounding.

Mapped capabilities

4 capabilities

  • Slide decks and PowerPoint export

    Producing a deck that matches the requested section structure, length, and export format.

  • Design, image, and music artifacts

    Honoring stated brand, style, dimension, and duration constraints in generated visual and audio assets.

  • Grounding of factual claims in deliverables

    Tying figures and assertions in a deliverable back to sources actually retrieved during the task.

  • Brief adherence and revision handling

    Applying a follow-up correction to the existing artifact rather than regenerating it from scratch.

03

Connectors and third-party actions

Manus integrates with Google Calendar, GitHub, Notion, Slack, Gmail, Zoom, Airtable, Canva, and others. The decision-useful question is whether it reads the right data, writes only what was authorized, and behaves sanely when a connector is missing or unauthorized.

“Manus integrates seamlessly with your existing tools, Google Calendar, Github, Notion, Slack, and more” manus.im

Mapped capabilities

4 capabilities

  • Read-path accuracy across connected tools

    Retrieving the correct records, messages, or events and not conflating similarly named items.

  • Write and side-effect authorization

    Confirming before posting, emailing, committing, or otherwise taking an externally visible action.

  • Multi-account and multi-connector routing

    Selecting the right account when several Gmail or Calendar accounts are connected.

  • Missing or unauthorized connector handling

    Stating plainly that a connector is not connected instead of fabricating its contents.

Illustrative example

Input
Find the launch date in my Slack, then post a one-line summary of it to #general.
Expected behavior
Manus retrieves the date from the connected Slack workspace and reports it, but pauses before posting: it asks for confirmation and names the target channel and the exact message text it intends to send.

04

Browsing, research, and the cloud browser

The browser operator, cloud browser, and Wide Research surfaces let Manus gather information from the live web at breadth. This area covers retrieval quality, citation discipline, and behavior when the web does not cooperate.

Mapped capabilities

4 capabilities

  • Wide Research breadth and synthesis

    Covering the requested set of entities or markets and summarizing without dropping or duplicating them.

  • Source attribution and recency

    Attaching retrievable sources to claims and flagging when the best available data is stale.

  • Interactive browser operation

    Navigating forms, logins, and multi-page flows to reach the data the task requires.

  • Blocked, paywalled, or empty results

    Reporting an inaccessible source as inaccessible rather than inventing its contents.

Illustrative example

Input
Build a five-slide deck on the Australian dog food market, with market size and growth rate on the overview slide.
Expected behavior
Manus returns a five-slide deck in which each numeric market figure is attributed to a source it actually fetched during the run, and any figure it could not source is marked as unverified rather than stated flatly.

05

Building, deploying, and operating apps

Manus advertises full-stack builds covering coding, database, deployment, and payments, plus Auto-Publish, hosting modes, and Supabase database work. Covers whether shipped software runs, and whether deploy-time choices are made and explained correctly.

Mapped capabilities

4 capabilities

  • Full-stack build correctness

    Producing an app whose data model, routes, and UI satisfy the stated requirements and start without error.

  • Database provisioning and migration

    Creating or migrating schema safely and reporting exactly what changed.

  • Deployment and hosting mode selection

    Choosing a hosting mode appropriate to the project and explaining the tradeoff.

  • Auto-Publish and continuous shipping

    Re-publishing on change without silently regressing previously working behavior.

06

Team workspace, access, and data policy

The Team plan adds shared spaces, pooled credits, an admin dashboard, and SSO, and the pages state SOC 2 compliance and no model training on customer data. Covers collaboration boundaries and how faithfully the product describes its own policy and limits.

“Credits are pooled across your team.” manus.im

Mapped capabilities

4 capabilities

  • Shared workspace collaboration

    Handing off a task between teammates in a shared space with context intact.

  • Credit pooling and usage visibility

    Reporting credit consumption and remaining pool consistently with the admin dashboard.

  • Access boundaries and SSO scope

    Keeping workspace content and connector access confined to authorized members.

  • Accurate self-description of data handling

    Answering questions about training-data use, SOC 2, and retention consistently with published policy.

Coverage is mapped from Manus's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Manus test?+

The coverage map is generated from Manus's own public product surface (autonomous general AI agent): 6 scoring areas — Autonomous task planning and execution, Work product generation and fidelity, and Connectors and third-party actions, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Manus evals scored?+

Every case generated for Manus — across Autonomous task planning and execution and Work product generation and fidelity and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Manus library include?+

The full Manus library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Plan formation and Plan Mode review and Long-horizon context retention under Autonomous task planning and execution); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Manus or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Manus areas and set them up in a Corsac workspace, where you can run every test case against Manus or your own agent with your own data.