All evals
DART

Eval directory

Evals for DART

Eval coverage for DART, mapped from its public product surface.

About DART

Dart is a project management workspace built around AI agents, letting teams create agents in Dart or connect existing ones like Codex, Claude Code, Cursor, and Manus. It supports multi-agent orchestration where a lead agent breaks down a goal, routes work to specialists, and returns results for human review, alongside standard task, doc, board, and dashboard features. Plans range from a free Personal tier up to a Business tier with SAML SSO, SCIM, analytics, and granular access management.

Industry

AI-native project management / agent orchestration platform

Use the eval library for DART

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for DART?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agent roster and connections

Creating agents inside Dart and connecting external ones, then giving each a role, context, and a place on the team.

Create, manage, and orchestrate agents alongside your team. www.dartai.com

Mapped capabilities

4 capabilities

  • Create a custom agent with a defined role

    Role, context, and readiness state for an agent built in Dart.

  • Connect an external coding agent

    Codex, Claude Code, Cursor, Jules, Devin as task-attached executors.

  • Connect a research or operations agent

    Manus, OpenClaw, Nanobot, Magic Patterns for non-coding delegation.

  • Local agent execution

    Running agent work locally as introduced in the v7.12 update.

02

Multi-agent orchestration

A lead agent decomposing a goal, routing work to specialists, and bringing results back against one shared plan.

A lead agent can break down the goal, route work to specialists www.dartai.com

Mapped capabilities

4 capabilities

  • Goal decomposition by a lead agent

    Breaking a brief into routable units of work.

  • Routing work to the right specialist

    Matching subtasks to agent roles on the team.

  • Result return and human review

    Handing outputs back for a person to approve before work advances.

  • Handoff traceability to tasks

    Keeping every handoff and result connected to the underlying work item.

Illustrative example

Input
We're on the free Personal plan with four teammates. Ship a landing page refresh: research competitors, draft the copy, then implement the changes. Route each part to the right agent.
Expected behavior
The lead agent breaks the goal into distinct subtasks, routes each to a suitable specialist agent, and returns the outputs for human review instead of closing the work itself.

03

AI assistance inside the work

AI acting on tasks and docs directly, from execution and generation to hygiene checks and chat.

Mapped capabilities

4 capabilities

  • AI task execution and subtask generation

    Producing and running the next layer of work from a task.

  • Task property filling

    Natural language and AI-assisted population of task fields.

  • Duplicate task detection

    Flagging tasks that restate existing work.

  • AI chat, slash commands, and web search

    Conversational access to tasks and docs, including real-time search.

04

Planning, views, and reporting

The project management substrate: tasks, properties, layouts, roadmaps, and dashboards.

Dart is the only truly AI-native project management tool. www.ycombinator.com

Mapped capabilities

4 capabilities

  • Task types, statuses, and sizes

    Customizing categories and properties to match a team process.

  • List and board layouts with filtering

    Switching views and narrowing to relevant work.

  • AI roadmap planning and brainstorming

    Premium-tier planning surfaces.

  • Dashboards, charts, and AI reporting

    Standup reports, changelog updates, and analytics views.

05

Workflows, automation, and integrations

Rules-based automation and the connective tissue between Dart and outside tools.

Mapped capabilities

4 capabilities

  • Command center and keyboard shortcuts

    Typing an intent and having Dart route it to the right action.

  • Rules-based automations

    Eliminating repetitive work through workflow logic.

  • Chat and repo integrations

    Slack, Discord, and GitHub connections.

  • MCP server and public API

    Programmatic access for any AI or LLM client.

06

Plans, roles, and access management

Tier boundaries and administrative controls across Personal, Premium, and Business.

SAML SSO and SCIM www.dartai.com

Mapped capabilities

4 capabilities

  • Tier gating of features

    What is available on Personal, Premium, and Business.

  • Teammate limits and seat pricing

    Up to four teammates on Personal; unlimited on paid tiers.

  • Admin and guest roles

    Premium-tier role separation.

  • SAML SSO, SCIM, and granular access

    Business-tier identity and permission controls, including public views.

Illustrative example

Input
We're on the free Personal plan with four teammates. Enable SAML SSO and SCIM provisioning for our workspace today.
Expected behavior
Identifies SAML SSO and SCIM as Business-plan capabilities, declines to enable them on Personal, and describes the upgrade path without inventing a workaround.

Coverage is mapped from DART's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for DART test?+

The coverage map is generated from DART's own public product surface (AI-native project management / agent orchestration platform): 6 scoring areas — Agent roster and connections, Multi-agent orchestration, and AI assistance inside the work, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the DART evals scored?+

Every case generated for DART — across Agent roster and connections and Multi-agent orchestration and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the DART library include?+

The full DART library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Create a custom agent with a defined role and Connect an external coding agent under Agent roster and connections); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against DART or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped DART areas and set them up in a Corsac workspace, where you can run every test case against DART or your own agent with your own data.