All evals
C

Eval directory

Evals for CodeRabbit

Eval coverage for CodeRabbit, mapped from its public product surface.

About CodeRabbit

CodeRabbit is an AI code review tool that automates pull request reviews with context-aware, line-by-line feedback on GitHub and GitLab. It also runs inside the terminal (CLI) and IDEs such as VS Code, Cursor, and Windsurf, reviewing uncommitted changes and offering one-click fixes. Plans range from a free tier to Pro, Pro Plus, and an Enterprise tier with RBAC, SSO, audit logging, and a self-hosting option.

Industry

AI code review platform

Use the eval library for CodeRabbit

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for CodeRabbit?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Pull request review on GitHub and GitLab

The hosted review surface: context-aware, line-by-line feedback delivered as bot comments on pull requests, plus the structured walkthrough, change-stack view, and merge-gating signals that accompany a review.

Mapped capabilities

4 capabilities

  • Line-by-line inline review comments

    Per-line feedback on changed hunks, including defects, code smells, refactors, and missed unit tests.

  • Walkthrough, summaries, and change stack

    PR summarization, per-file and per-change summaries, ordered change stack, and files/overview navigation.

  • Pre-merge checks and merge blockers

    Built-in pre-merge checks, custom checks on Pro Plus, pass/fail state, and not-mergeable blocker reporting.

  • Impact framing: blast radius and architecture impact

    Review-level signals describing the reach and architectural effect of a change set.

Illustrative example

Input
@coderabbitai why can't I merge this PR yet?
Expected behavior
Names the failed pre-merge checks and the outstanding blockers shown in the current review state, and does not describe the pull request as mergeable or ready.

02

In-flow review in CLI and IDE

Review that runs where code is written rather than on a pull request: the terminal CLI and the VS Code, Cursor, and Windsurf extensions, acting as a quality gate for both human and agent-authored changes.

Cut code review time & bugs in half, instantly. www.coderabbit.ai

Mapped capabilities

4 capabilities

  • Review of uncommitted staged and unstaged changes

    Immediate feedback on local working-tree changes before a PR is raised.

  • One-click fixes

    Applying suggested review fixes back into the codebase from the CLI or editor.

  • Install and platform support

    curl install script for macOS, Linux, and Windows (WSL); extension install across VS Code, Cursor, and Windsurf.

  • Free-tier reviews and rate limits

    No-cost CLI and IDE reviews subject to rate limits, with usage-based credits as the unrestricted path.

03

Agentic chat, commands, and follow-through

Conversational and agentic control of the reviewer: invoking commands on a PR, chatting about a change, and the automated actions CodeRabbit takes to finish or follow up on work.

CodeRabbit is designed to work with all programming languages www.coderabbit.ai

Mapped capabilities

4 capabilities

  • Command invocation and help

    @coderabbitai commands on a pull request, including discovery of available commands via help.

  • Agentic chat with CodeRabbit

    Conversational follow-up about a review, change, or comment.

  • Finishing touches

    Docstring generation and autofix on Pro; unit test generation, simplify, and merge conflict resolution on Pro Plus.

  • Issue planner and post-merge actions

    Pro Plus issue planning, and post-merge follow-through driven by the reviewed pull request.

04

Toolchain and workflow integrations

How CodeRabbit connects to the surrounding development stack — static analysis tooling, issue trackers, external context via MCP, and reporting back to the team.

Mapped capabilities

4 capabilities

  • Linters and SAST tool support

    Incorporating linter and static application security testing results into reviews.

  • Jira and Linear integrations

    Issue tracker linkage from reviews and planning.

  • MCP connections and linked repository analysis

    Pulling additional context from connected sources and related repositories.

  • Analytics dashboards and customizable reports

    Product analytics on review activity and configurable reporting for teams.

05

Plans, entitlements, and billing

Tier boundaries and purchase mechanics: what Free, Pro, Pro Plus, and Enterprise each include, how trials and billing cadence work, and how usage-based credits extend limits.

CodeRabbit offers a free 14-day Pro Plus trial with no credit card required. www.coderabbit.ai

Mapped capabilities

4 capabilities

  • Tier feature boundaries

    Which capabilities belong to Free, Pro, Pro Plus, and Enterprise, and where each line is drawn.

  • Trial and signup

    14-day Pro Plus trial with no credit card, and two-click install via GitHub or GitLab.

  • Billing cadence and per-seat pricing

    Monthly versus annual billing, annual savings, and per-user pricing.

  • Usage-based credit add-on

    Unrestricted CLI and PR reviews via one-time or subscription credit purchases, managed from the dashboard.

Illustrative example

Input
On the free plan, do I get line-by-line review comments on my pull requests, or only summaries?
Expected behavior
States that the Free plan covers PR summarization across unlimited public and private repositories, notes the 14-day Pro Plus trial needs no credit card, and flags that free CLI and IDE reviews are rate limited.

06

Enterprise security, privacy, and governance

The controls an Enterprise buyer evaluates: access management, deployment topology, compliance posture, and control over how code and review data are retained.

Mapped capabilities

4 capabilities

  • Access control and identity

    Custom RBAC, SSO, and multi-org support.

  • Auditability and API access

    Audit logging and programmatic API access on Enterprise.

  • Deployment topology

    Self-hosting option, EU SaaS deployment, and custom setup such as ALB.

  • Compliance and data handling

    SOC 2 Type II, GDPR compliance, encryption practices, and opting out of data storage.

Coverage is mapped from CodeRabbit's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for CodeRabbit test?+

The coverage map is generated from CodeRabbit's own public product surface (AI code review platform): 6 scoring areas — Pull request review on GitHub and GitLab, In-flow review in CLI and IDE, and Agentic chat, commands, and follow-through, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the CodeRabbit evals scored?+

Every case generated for CodeRabbit — across Pull request review on GitHub and GitLab and In-flow review in CLI and IDE and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the CodeRabbit library include?+

The full CodeRabbit library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Line-by-line inline review comments and Walkthrough, summaries, and change stack under Pull request review on GitHub and GitLab); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against CodeRabbit or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped CodeRabbit areas and set them up in a Corsac workspace, where you can run every test case against CodeRabbit or your own agent with your own data.