All evals
Q

Eval directory

Evals for Qodo

Eval coverage for Qodo, mapped from its public product surface.

About Qodo

Qodo is an AI code review and governance platform for engineering teams that runs multi-agent reviews on pull requests with full codebase context. It centralizes coding standards through a rules system, surfaces cross-repo breaking changes and requirement gaps, and provides dashboards for tracking findings across repos and teams. It integrates with Git providers and IDEs (including agentic IDEs like Kiro), with a paid Pro Team tier and an enterprise tier offering SSO/SAML, BYOK, and self-hosted deployment.

Industry

AI code review and governance platform

Use the eval library for Qodo

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Qodo?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agentic PR Code Review

Specialized agents run on every pull request with full codebase context, surfacing real bugs, logic gaps, rule violations, and requirement gaps as actionable suggestions.

Qodo empowers engineering teams to standardize code quality and code review to move fast with AI. www.qodo.ai

Mapped capabilities

4 capabilities

  • Issue and logic-gap detection

    Identifies critical defects in the diff rather than surface style commentary.

  • Requirement gap detection

    Flags where the change does not fulfill the stated intent of the PR.

  • Review depth selection

    Applies review effort modes / adaptive depth appropriate to the PR.

  • Actionable suggestion quality

    Findings are specific, located, and paired with a proposed fix.

02

Rules and Standards Governance

A living rules system where coding standards are defined, edited, and enforced in one place, kept consistent and up to date as code and teams change.

Mapped capabilities

4 capabilities

  • Rule authoring and editing

    Define and update standards centrally with no rule-count limit.

  • Deterministic rule enforcement

    The same rule resolves consistently for developers, reviewers, and AI agents.

  • Rule Miner

    Codifies conventions implied by existing review history into explicit rules.

  • Self-learning rules

    Advanced self-learning behavior offered on enterprise tiers.

Illustrative example

Input
Review this PR. It adds an endpoint that logs the full request body. Our rules system contains a rule: never log request payloads that may contain user PII.
Expected behavior
The review flags the logging call as a violation of the PII rule, cites that specific rule by name, points to the offending line, and proposes a redacted-logging fix instead of returning a clean verdict.

03

Cross-Repo Context

Reasoning about the system rather than the diff: repo relationships, contract verification, and changes whose blast radius lands outside the repo under review.

Specialized agents run on every pull request, surfacing real bugs, rule violations, and requirement gaps with full codebase context. www.qodo.ai

Mapped capabilities

4 capabilities

  • Breaking change detection

    Surfaces consumers broken by a change to a shared interface.

  • Contract verification across repos

    Checks producer and consumer expectations stay aligned.

  • Dependency conflict surfacing

    Flags conflicting dependency expectations across repos.

  • Repo relationship mapping

    Represents which repos depend on which.

Illustrative example

Input
Review this PR in service-a that renames a field in a shared API response schema. Another repo, service-b, reads that field.
Expected behavior
The review identifies the rename as a cross-repo breaking change and names service-b as an affected consumer, rather than treating the change as a self-contained diff-level edit.

04

Shift-Left IDE and Agent Workflows

Review that runs before the PR exists — inside the IDE, inside the developer's coding agent, and in the CLI — so findings and fixes surface at coding time.

High-precision multi-agent code review across the SDLC, with full codebase context. www.qodo.ai

Mapped capabilities

4 capabilities

  • Pre-PR review skills

    Review skills run inside the developer's agent earlier in the lifecycle.

  • Agentic IDE integration

    Runs inside agentic IDEs such as Kiro alongside the workspace agents.

  • Local fix and validation loop

    Developers resolve flagged code in context and validate the change.

  • CLI and SDLC automation

    Quality agents invoked outside the PR gate.

05

Governance Visibility and Analytics

One portal for tracking findings, resolution rates, and open issues across every repo, pull request, team, and AI tool, with audit trails for enterprise oversight.

Track critical findings, resolution rates and open issues across every repo and pull request in one view www.qodo.ai

Mapped capabilities

4 capabilities

  • Cross-repo findings view

    Critical findings and open issues aggregated in a single view.

  • Resolution rate tracking

    Reports whether surfaced findings were actually addressed.

  • Team and repo breakdowns

    Segments governance signal by owning team and repository.

  • Audit logs

    Enterprise-tier record of governance and access activity.

06

Enterprise Deployment and Plan Controls

The commercial and deployment envelope: Git provider coverage, identity and key management, hosting model, and the credit-based Pro Team billing controls.

Mapped capabilities

4 capabilities

  • SSO/SAML and access control

    Enterprise identity integration for org-wide rollout.

  • BYOK and data retention

    Bring-your-own LLM keys and strict retention commitments.

  • Deployment topology

    Single-tenant SaaS, on-premises, and air-gapped options.

  • Credit packs and overage caps

    Pooled credits, pack switching, and customer-set monthly overage cap.

Coverage is mapped from Qodo's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Qodo test?+

The coverage map is generated from Qodo's own public product surface (AI code review and governance platform): 6 scoring areas — Agentic PR Code Review, Rules and Standards Governance, and Cross-Repo Context, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Qodo evals scored?+

Every case generated for Qodo — across Agentic PR Code Review and Rules and Standards Governance and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Qodo library include?+

The full Qodo library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Issue and logic-gap detection and Requirement gap detection under Agentic PR Code Review); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Qodo or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Qodo areas and set them up in a Corsac workspace, where you can run every test case against Qodo or your own agent with your own data.