
Eval directory
Evals for Windsurf
8 evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Windsurf AI products.
About Windsurf
Windsurf (by Codeium) is an agentic AI IDE. Its Cascade agent does multi-file, plan-and-act coding with terminal access, alongside predictive Tab / Supercomplete completions, local codebase indexing and @-mentions, persistent Memories and .windsurfrules, Flows that keep the AI and human in shared state, MCP integrations, and a multi-model picker.
How complete this published benchmark library is across datasets, metrics, rubrics, use-case maps, and pack context. This is library coverage, not an agent performance score.
Test datasets
8/8 packs
Scoring metrics
0/8 packs
Judge rubrics
8/8 packs
Use-case maps
0/8 packs
Pack context
8/8 packs
Available eval packs for Windsurf
8 packs ready to run.
Cascade Agent
Evaluates Windsurf's Cascade Agent across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Code Editor eval coverage.
Context And Indexing
Evaluates Windsurf's Context & Indexing across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Code Editor eval coverage.
Flows And Terminal
Evaluates Windsurf's Flows & Terminal across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Code Editor eval coverage.
Mcp And Integrations
Evaluates Windsurf's MCP & Integrations across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Code Editor eval coverage.
Memories And Rules
Evaluates Windsurf's Memories & Rules across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Code Editor eval coverage.
Models And Credits
Evaluates Windsurf's Models & Credits across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Code Editor eval coverage.
Safety Privacy And Governance
PII LeakageEvaluates Windsurf's Safety, Privacy & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Code Editor eval coverage.
Tab Autocomplete And Supercomplete
Evaluates Windsurf's Tab / Autocomplete / Supercomplete across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Code Editor eval coverage.
Why eval Windsurf AI
Windsurf's AI features ship behind brand promises about accuracy, safety, and reliability. Buyers and integrators need to know those promises hold up under adversarial prompts, edge-case workflows, and the long tail of real customer inputs — not just the demo path.
The Corsac eval library for Windsurf measures four dimensions teams care about most when deploying code assistant agents:
- Adversarial robustness — does the agent resist prompt injection, jailbreaks, and social-engineering attempts?
- Workflow quality— does it complete the task buyers were shown in the demo, on inputs that don't look like the demo?
- Safety gates — does it escalate or refuse when it should, and only then?
- Operator quality — does it preserve analyst trust by surfacing the right context at the right time?
Every eval pack above is hand-authored against Windsurf's public product surface and runnable in Corsac with your own data.