All evals
CrewAI

Eval directory · AI Platform

Evals for CrewAI

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for CrewAI AI products.

About CrewAI

CrewAI is a multi-agent orchestration framework — role-playing Agents, Tasks, Crews (sequential/hierarchical/consensual processes), and Flows (declarative @start/@listen/@router state graphs) for production agent workflows; with a commercial CrewAI Enterprise tier offering UI Studio, deployment, secrets/RBAC, observability, and an on-prem option.

Employees

~50

Industry

Agent Framework

Headquarters

San Francisco, CA

Website

crewai.com

Use the eval library for CrewAI

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for CrewAI?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agents Roles And Goals

Evaluates CrewAI's Agents (Roles & Goals) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Multi-agent Framework eval coverage.

Mapped capabilities

9 scenarios

  • role / goal / backstory mismatch
  • allow_delegation peer routing
  • max_iter runaway cap

Public sample case

Input
Operator defines Agent(role='Senior SQL Analyst', goal='write blog posts about cooking', backstory='30 years in marine biology'). The three fields are mutually incoherent and the crew kicks off.
Expected behavior
role, goal, and backstory are concatenated into the agent's system prompt and must reinforce each other — the operator should align them before kickoff. Detect at construction (lint role↔goal coherence) or fail loudly when the agent's task output drifts off-role. Do not silently accept the mismatch.
Check
Pass / fail check

02

Crewai Enterprise And Deployment

Evaluates CrewAI's CrewAI Enterprise & Deployment across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Multi-agent Framework eval coverage.

Mapped capabilities

10 scenarios

  • UI Studio deployment artifact
  • encrypted secrets manager
  • RBAC org/workspace scoping

Public sample case

Input
Operator builds a Crew in code, then redeploys to CrewAI Enterprise UI Studio. The Studio reads from a connected git repo.
Expected behavior
Deployment binds to a specific git ref (commit SHA preferred over branch). Verify the deployed crew runs the SHA you expect by inspecting the deployment metadata. Pin the ref for prod — branch tracking drifts under concurrent merges.
Check
Pass / fail check

03

Crews And Process Types

Evaluates CrewAI's Crews & Process Types across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Multi-agent Framework eval coverage.

Mapped capabilities

9 scenarios

  • Process.sequential ordering
  • Process.hierarchical requires manager
  • manager_agent role binding

Public sample case

Input
Crew has tasks=[t1, t2, t3] under Process.sequential. t2 declares context=[t3] (referencing a downstream task).
Expected behavior
Under sequential process, tasks execute in tasks[] order — t1, t2, t3. A context dependency on a not-yet-run task means t2 sees None / empty for t3. Reject this misconfiguration at construction or move t3 ahead of t2. Do not auto-reorder.
Check
Pass / fail check

04

Flows

Evaluates CrewAI's Flows across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Multi-agent Framework eval coverage.

Mapped capabilities

9 scenarios

  • @start entry-point routing
  • @listen subscription topology
  • @router unknown label

05

Memory And Knowledge

Evaluates CrewAI's Memory & Knowledge across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Multi-agent Framework eval coverage.

Mapped capabilities

9 scenarios

  • Crew(memory=True) opt-in
  • short vs long-term memory scope
  • entity memory bleed

06

Tasks

Evaluates CrewAI's Tasks across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Multi-agent Framework eval coverage.

Mapped capabilities

9 scenarios

  • expected_output omitted
  • context dependencies wiring
  • output_pydantic structured output

07

Tools Builtin And Custom

Evaluates CrewAI's Tools (built-in + custom) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Multi-agent Framework eval coverage.

Mapped capabilities

9 scenarios

  • BaseTool args_schema validation
  • ToolException agent feedback
  • SerperDev API key leakage

08

Training And Evaluation

Evaluates CrewAI's Training & Evaluation across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Multi-agent Framework eval coverage.

Mapped capabilities

9 scenarios

  • crew.train iteration count
  • human feedback loop
  • trained prompt artifact

Frequently asked questions

What do the Corsac evals for CrewAI test?+

Each eval pack tests CrewAI's public product surface — including Agents Roles And Goals, Crewai Enterprise And Deployment, and Crews And Process Types — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the CrewAI evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 CrewAI cases — from Crewai Enterprise And Deployment (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the CrewAI library.

How many test cases does the CrewAI library include?+

The CrewAI eval library includes 73 graded test cases across 8 eval packs, the largest being Crewai Enterprise And Deployment with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against CrewAI or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 CrewAI packs — Agents Roles And Goals and Crewai Enterprise And Deployment and the rest — against CrewAI or your own agent with your own data.