All evals
C

Eval directory

Evals for Codacy

Eval coverage for Codacy, mapped from its public product surface.

About Codacy

Codacy is a code quality and security platform that scans code across AI agents, IDEs, Git, containers, and runtime. It embeds automated reviews, guardrails, and auto-fixes so AI-generated code meets an organization's quality and security standards before merge. It integrates with GitHub, GitLab, and Bitbucket and adds pull request gating, coverage policies, and compliance evidence.

Employees

57

Industry

code quality & security platform for AI-assisted engineering

Use the eval library for Codacy

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Codacy?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

AI Agent & IDE Guardrails

Pre-commit and in-prompt enforcement: scanning as code is generated or typed, auto-fixing findings, and handing issues back to the agent so generated code meets standards before a pull request opens.

Embed security checks and auto-fixes on every prompt www.codacy.com

Mapped capabilities

4 capabilities

  • Scan-as-you-type findings in the editor

    Surfacing SAST, SCA, secrets, and quality violations inside VSCode, JetBrains, and Cursor.

  • Auto-fix and agent handoff

    Proposing fixes and routing issues back to the agent for correction before commit.

  • AI policy violations and unapproved model calls

    Flagging generated code that breaks organizational AI policy or calls unapproved models.

  • Guardrail behavior across agent and IDE entry points

    Consistency of the same checks whether invoked from an agent prompt or the IDE extension.

02

Pull Request Review & Merge Gating

The Git-stage workflow: automated review comments on pull requests, merge gates tied to standards and coverage, and noise control so teams merge quickly without shipping new defects.

Shared Coding Standards across 49 languages www.codacy.com

Mapped capabilities

4 capabilities

  • AI Reviewer comments and fix suggestions

    Low-noise, actionable PR feedback with ready-to-commit suggestions and PR summaries.

  • Merge gates against coding standards

    Blocking or allowing merges based on configured quality and security standards.

  • Coverage reports and merge policies

    Gating on unit test coverage and untested-code policies.

  • False positive detection and triage

    Suppressing or marking findings the platform identifies as false positives.

Illustrative example

Input
Our PR has one unresolved secret-scanning finding and the merge gate is on. Can we merge, and what are our options to unblock it?
Expected behavior
States the merge is blocked while the finding is unresolved, and points to resolving or explicitly handling the finding as the path forward. Does not invent a severity threshold, bypass flag, or approval role that the documented behavior does not describe.

03

Security Scanning Coverage

The breadth of security analysis the platform advertises across code, dependencies, infrastructure definitions, images, and running applications, and which scan type applies to a given concern.

Make every line of AI generated code follow your quality & security standards by default. www.codacy.com

Mapped capabilities

4 capabilities

  • SAST and secret scanning

    Static findings including SQL injection classes and committed secrets.

  • Dependency and package risk (SCA)

    Insecure dependencies, malicious package detection, and license scanning.

  • Infrastructure-as-code scanning

    IaC misconfiguration findings at the Git stage.

  • Container image and runtime scanning

    Container CVEs before deployment, plus DAST and pen-testing for runtime apps and API endpoints.

04

Coding Standards & Quality Analytics

Shared, org-wide definitions of acceptable code and the reporting that shows whether teams are converging on them over time.

Mapped capabilities

4 capabilities

  • Shared coding standards across languages

    Applying one standard across the supported language set and multiple repositories.

  • Maintainability findings

    Complex code, error-prone code, unused code, and code duplication.

  • Trends across teams and projects

    Reporting movement in quality metrics over time.

  • Custom rules

    Organization-defined rules beyond the built-in rule set.

05

Git Provider Integration & Onboarding

Connecting repositories and getting to a first result: provider-based login, self-serve scanning, and the enterprise variants that change how integration is set up.

Mapped capabilities

4 capabilities

  • GitHub, GitLab, and Bitbucket connection

    Provider login and repository connection paths.

  • Self-serve first scan

    One-click signup and scanning a first repository without a sales call.

  • Repository and project scope limits

    Private repo counts and unlimited-project boundaries by plan.

  • GitHub Enterprise and data residency

    Enterprise-hosted Git and data residency handling.

06

Governance, Compliance & Administration

Administrative and evidentiary surfaces: proving to auditors what was enforced, controlling access, and understanding AI-related risk across the codebase.

Mapped capabilities

4 capabilities

  • Compliance evidence generation

    Producing evidence that quality and security controls ran and gated merges.

  • AI Inventory and AI Risk Hub

    Visibility into AI usage and associated risk across repositories.

  • SSO/SAML and audit logs

    Enterprise access control and administrative traceability.

  • Plan and entitlement boundaries

    Which capabilities are available on free, paid, and enterprise tiers.

Illustrative example

Input
We're on the free plan for open-source repos. Is container image scanning and DAST included, or do we need to upgrade?
Expected behavior
Says container image scanning and DAST are not part of the free tier and belong to the higher enterprise tier, then points to the demo or sales path. Does not describe them as available on free or open-source plans.

Coverage is mapped from Codacy's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Codacy test?+

The coverage map is generated from Codacy's own public product surface (code quality & security platform for AI-assisted engineering): 6 scoring areas — AI Agent & IDE Guardrails, Pull Request Review & Merge Gating, and Security Scanning Coverage, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Codacy evals scored?+

Every case generated for Codacy — across AI Agent & IDE Guardrails and Pull Request Review & Merge Gating and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Codacy library include?+

The full Codacy library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Scan-as-you-type findings in the editor and Auto-fix and agent handoff under AI Agent & IDE Guardrails); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Codacy or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Codacy areas and set them up in a Corsac workspace, where you can run every test case against Codacy or your own agent with your own data.