All evals
Cursor

Eval directory · Code Assistant

Evals for Cursor

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Cursor AI products.

About Cursor

Cursor is an AI code editor built on VS Code: predictive Tab completion, inline edits, and an agent mode that plans and executes multi-file changes with terminal access, codebase indexing, project rules, and MCP integration.

Employees

~200

Industry

AI Code Editor

Headquarters

San Francisco, CA

Website

cursor.com

Use the eval library for Cursor

All 51 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Cursor?

9 areas · 51 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Codebase Indexing

Evaluates Cursor's Codebase Indexing across 6 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Code Editor eval coverage.

Mapped capabilities

6 scenarios

  • @codebase retrieval grounding
  • @file symbol precision
  • stale index after branch switch

02

Composer Agent

Evaluates Cursor's Composer & Agent Mode across 7 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Code Editor eval coverage.

Mapped capabilities

7 scenarios

  • multi-file plan before write
  • terminal command approval gate
  • .cursorignore write scope

03

Inline Edit

Evaluates Cursor's Inline Edit (Cmd-K) across 6 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Code Editor eval coverage.

Mapped capabilities

6 scenarios

  • selection-scoped intent fidelity
  • diff preview before apply
  • no out-of-scope file mutations

04

Mcp Integration

Evaluates Cursor's MCP Integration across 6 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Code Editor eval coverage.

Mapped capabilities

6 scenarios

  • mcp.json server connection
  • tool approval before invoke
  • MCP error and timeout handling

05

Model Selection

Evaluates Cursor's Model Selection & Routing across 5 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Code Editor eval coverage.

Mapped capabilities

5 scenarios

  • manual model picker override
  • auto mode routing
  • max/thinking budget modes

06

Privacy Edit Safety

Evaluates Cursor's Privacy & Edit Safety across 6 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Code Editor eval coverage.

Mapped capabilities

6 scenarios

  • Privacy Mode ZDR enforcement
  • dotfile protection .env
  • YOLO auto-run disabled

07

Project Rules

Evaluates Cursor's Project Rules across 6 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Code Editor eval coverage.

Mapped capabilities

6 scenarios

  • .cursor/rules mdc frontmatter
  • AGENTS.md vs rules precedence
  • alwaysApply vs glob matching

08

Tab Completion

Evaluates Cursor's Tab Completion across 6 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Code Editor eval coverage.

Mapped capabilities

6 scenarios

  • multi-line prediction acceptance
  • language syntax and import context
  • partial accept vs full tab accept

09

Completion Smoke V1

Smoke test for AI coding / IDE assistant completion correctness.

Mapped capabilities

3 scenarios

  • Completion Correctness
  • Style and Maintainability
  • Execution Safety

Example criterion: Cursor generates correct, maintainable code completions that satisfy task intent without unsafe patterns.

Frequently asked questions

What do the Corsac evals for Cursor test?+

Each eval pack tests Cursor's public product surface — including Codebase Indexing, Composer Agent, Inline Edit — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Cursor evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Cursor library include?+

The Cursor eval library includes 51 graded test cases across 9 eval packs. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Cursor or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against Cursor or your own agent with your own data.