
Eval directory
Evals for Cognition
8 evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Cognition AI products.
About Cognition
Cognition builds Devin, an autonomous AI software engineer that plans, writes, debugs, and ships code in a sandboxed cloud environment with terminal, browser, and editor access, session continuity, and human-in-the-loop review.
How complete this published benchmark library is across datasets, metrics, rubrics, use-case maps, and pack context. This is library coverage, not an agent performance score.
Test datasets
8/8 packs
Scoring metrics
0/8 packs
Judge rubrics
8/8 packs
Use-case maps
0/8 packs
Pack context
8/8 packs
Available eval packs for Cognition
8 packs ready to run.
Code Generation And Refactoring
Code CheckerEvaluates Cognition's Code Generation & Refactoring across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.
Devin Sessions And Planning
Evaluates Cognition's Devin Sessions & Planning across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.
Human In The Loop And Review
Evaluates Cognition's Human-in-the-loop & Review across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.
Knowledge And Memory
Knowledge RetentionEvaluates Cognition's Knowledge & Memory across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.
Repo Codebase Operations
Evaluates Cognition's Repo / Codebase Operations across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.
Safety Secrets And Governance
Evaluates Cognition's Safety, Secrets & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.
Sandbox Environment
Evaluates Cognition's Sandbox Environment across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.
Tool Use And Function Orchestration
Tool SelectionEvaluates Cognition's Tool Use & Function Orchestration across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.
Why eval Cognition AI
Cognition's AI features ship behind brand promises about accuracy, safety, and reliability. Buyers and integrators need to know those promises hold up under adversarial prompts, edge-case workflows, and the long tail of real customer inputs — not just the demo path.
The Corsac eval library for Cognition measures four dimensions teams care about most when deploying code assistant agents:
- Adversarial robustness — does the agent resist prompt injection, jailbreaks, and social-engineering attempts?
- Workflow quality— does it complete the task buyers were shown in the demo, on inputs that don't look like the demo?
- Safety gates — does it escalate or refuse when it should, and only then?
- Operator quality — does it preserve analyst trust by surfacing the right context at the right time?
Every eval pack above is hand-authored against Cognition's public product surface and runnable in Corsac with your own data.