All evals
Cognition

Eval directory · Code Assistant

Evals for Cognition

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Cognition AI products.

About Cognition

Cognition builds Devin, an autonomous AI software engineer that plans, writes, debugs, and ships code in a sandboxed cloud environment with terminal, browser, and editor access, session continuity, and human-in-the-loop review.

Employees

~200

Industry

Autonomous Coding Agent

Headquarters

San Francisco, CA

Use the eval library for Cognition

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Cognition?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Code Generation And Refactoring

Evaluates Cognition's Code Generation & Refactoring across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.

Mapped capabilities

9 scenarios

  • multi-file change atomicity
  • type safety on TypeScript edits
  • idempotent edits on str_replace

Public sample case

Input
Task: 'rename Order.total to Order.totalCents and convert dollars→cents at all call sites'. Spans 12 files.
Expected behavior
Land the rename and the conversion together (single PR) so no intermediate commit leaves the repo inconsistent. Run the typechecker + tests after each save group; do not push partial change. If the operator wants incremental review, stage commits within the same PR rather than separate PRs.
Check
Pass / fail check

02

Devin Sessions And Planning

Evaluates Cognition's Devin Sessions & Planning across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.

Mapped capabilities

9 scenarios

  • POST /v1/sessions create with snapshot_id
  • plan-vs-execute split visibility
  • POST /v1/sessions/{id}/message follow-up

Public sample case

Input
Operator creates a Devin session via POST /v1/sessions passing snapshot_id from a prior session to resume on the same VM state.
Expected behavior
Pass snapshot_id only when the prior session was NOT terminated — terminate is irreversible and invalidates the snapshot. On 4xx referencing an invalid snapshot, surface to operator with a 'create fresh session' fallback, do not loop the retry. Record the new session_id and persist the snapshot lin…
Check
Pass / fail check

03

Human In The Loop And Review

Evaluates Cognition's Human-in-the-loop & Review across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.

Mapped capabilities

9 scenarios

  • clarification question quality
  • ACU burn-rate alert to operator
  • blocker reporting with evidence

Public sample case

Input
Task is ambiguous: 'fix the failing tests'. Multiple test files are failing with unrelated causes.
Expected behavior
Ask a specific clarification with concrete options: 'Failing tests cluster into 3 groups — (A) auth helper rename, (B) DB migration drift, (C) flake in browser tests. Which should I prioritize?' Surface evidence (test names, last failure timestamps). Do not ask vague 'what do you want me to do?'
Check
Pass / fail check

04

Knowledge And Memory

Evaluates Cognition's Knowledge & Memory across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.

Mapped capabilities

9 scenarios

  • POST /v3/.../knowledge-notes scoping
  • Knowledge vs live-repo conflict
  • DeepWiki citation fidelity

05

Repo Codebase Operations

Evaluates Cognition's Repo / Codebase Operations across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.

Mapped capabilities

9 scenarios

  • GitHub install with selectable repo permissions
  • branch hygiene per task
  • PR title and body discipline

06

Safety Secrets And Governance

Evaluates Cognition's Safety, Secrets & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.

Mapped capabilities

10 scenarios

  • secret never echoed to terminal
  • prod-write boundary
  • prompt-injection from repo content

07

Sandbox Environment

Evaluates Cognition's Sandbox Environment across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.

Mapped capabilities

9 scenarios

  • terminal command capture
  • interactive browser handoff
  • VM ephemeral filesystem on snapshot

08

Tool Use And Function Orchestration

Evaluates Cognition's Tool Use & Function Orchestration across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.

Mapped capabilities

9 scenarios

  • terminal command shell selection
  • MCP server integration
  • browser action retry on transient nav

Frequently asked questions

What do the Corsac evals for Cognition test?+

Each eval pack tests Cognition's public product surface — including Code Generation And Refactoring, Devin Sessions And Planning, and Human In The Loop And Review — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Cognition evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Cognition cases — from Safety Secrets And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Cognition library.

How many test cases does the Cognition library include?+

The Cognition eval library includes 73 graded test cases across 8 eval packs, the largest being Safety Secrets And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Cognition or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Cognition packs — Code Generation And Refactoring and Devin Sessions And Planning and the rest — against Cognition or your own agent with your own data.