
Eval directory
Evals for Sourcegraph
8 evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Sourcegraph AI products.
About Sourcegraph
Sourcegraph is a code intelligence and AI coding platform: universal code search, precise code navigation, Cody chat grounded in your codebase, cross-repo batch changes, and the Amp autonomous agent — deployed across large enterprise codebases.
How complete this published benchmark library is across datasets, metrics, rubrics, use-case maps, and pack context. This is library coverage, not an agent performance score.
Test datasets
8/8 packs
Scoring metrics
0/8 packs
Judge rubrics
8/8 packs
Use-case maps
0/8 packs
Pack context
8/8 packs
Available eval packs for Sourcegraph
8 packs ready to run.
Amp Autonomous Agent
Evaluates Sourcegraph's Amp Autonomous Agent across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Code Intelligence eval coverage.
Batch Changes
Evaluates Sourcegraph's Batch Changes across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Code Intelligence eval coverage.
Code Insights And Ownership
Evaluates Sourcegraph's Code Insights & Ownership across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Code Intelligence eval coverage.
Cody Autocomplete And Inline Edit
Evaluates Sourcegraph's Cody Autocomplete & Inline Edit across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Code Intelligence eval coverage.
Cody Chat And Context
Evaluates Sourcegraph's Cody Chat & Context across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Code Intelligence eval coverage.
Deployment Auth And Governance
Evaluates Sourcegraph's Deployment, Auth & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Code Intelligence eval coverage.
Precise Code Navigation
Evaluates Sourcegraph's Precise Code Navigation across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Code Intelligence eval coverage.
Universal Code Search
Evaluates Sourcegraph's Universal Code Search across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Code Intelligence eval coverage.
Why eval Sourcegraph AI
Sourcegraph's AI features ship behind brand promises about accuracy, safety, and reliability. Buyers and integrators need to know those promises hold up under adversarial prompts, edge-case workflows, and the long tail of real customer inputs — not just the demo path.
The Corsac eval library for Sourcegraph measures four dimensions teams care about most when deploying code assistant agents:
- Adversarial robustness — does the agent resist prompt injection, jailbreaks, and social-engineering attempts?
- Workflow quality— does it complete the task buyers were shown in the demo, on inputs that don't look like the demo?
- Safety gates — does it escalate or refuse when it should, and only then?
- Operator quality — does it preserve analyst trust by surfacing the right context at the right time?
Every eval pack above is hand-authored against Sourcegraph's public product surface and runnable in Corsac with your own data.