All evals
Sonar

Eval directory · Security Operations

Evals for Sonar

Eval coverage for Sonar, mapped from its public product surface.

About Sonar

Sonar is a code verification platform for AI-generated and human-written code, built around a Guide → Verify → Solve loop. Its offerings include SonarQube (Server, Cloud, Community) with an Advanced Security add-on for dependency/SCA analysis, Gitar for agentic AI code review that applies real fixes, Sonar Vortex for injecting project context and real-time verification into agent coding loops, and SonarSweep (early access) for improving coding-LLM training data. It integrates with CI/CD pipelines and is sold via Team and Enterprise SonarQube plans plus Gitar Core/Pro per-user tiers.

Industry

AI code review and verification platform

Use the eval library for Sonar

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Sonar?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Code Verification & Analysis (SonarQube)

Core static and algorithmic analysis of AI-generated and human-written code across reliability, security, quality, and maintainability, with findings that are consistent, explainable, repeatable, and auditable.

Mapped capabilities

4 capabilities

  • Static analysis across supported languages

    30+ languages on Team; 40+ on Enterprise including ABAP, COBOL, Apex. Bug, vulnerability, and security hotspot detection.

  • Quality gates and technical debt metrics

    Maintainability, reliability, and technical debt tracking; quality gate at sandbox or pipeline exit.

  • Branch, pull request, and merge scanning

    Automatic scanning of all branches, PRs, and merges as code arrives; pull request analysis.

  • AI CodeFix and remediation suggestions

    AI-driven, one-click fix suggestions for bugs, vulnerabilities, and quality issues.

02

Supply Chain & Advanced Security

The Advanced Security add-on and security-reporting capabilities that extend SonarQube beyond first-party code into dependencies and compliance standards.

injects the right project context and constraints before the first line of code www.sonarsource.com

Mapped capabilities

4 capabilities

  • Dependency and SCA assessment

    Software supply chain and dependency analysis to prevent agents from introducing third-party vulnerabilities.

  • Secrets detection

    Included from the Team plan onward.

  • Compliance standard reporting

    OWASP, CWE, PCI DSS, and MISRA C++:2023 reporting on Enterprise.

  • Advanced security reports and audit logs

    Enterprise-tier reporting and auditability of findings.

03

Agentic Code Review & Fixes (Gitar)

Agentic review that produces working fixes rather than comments, iterating against CI until the pipeline passes, and acting as an interactive participant on pull and merge requests.

Gitar understands code context, generates working fixes, and iterates until your CI pipeline passes. www.sonarsource.com

Mapped capabilities

4 capabilities

  • PR review, summaries, and interactive agent

    Customizable reviews, automatic PR summaries, answering questions and making revisions in-thread.

  • CI failure analysis and de-duplication

    GitHub Actions and GitLab Pipelines on Core; CircleCI, Buildkite, Bitrise on Pro. Flakiness detection and root-cause analysis.

  • Fix generation and auto-apply

    Fixes via comments on Core; auto-apply until the PR is green on Pro.

  • Merge gating

    Auto-approve and merge blocking on issues, Pro tier only.

04

Agent Coding Loop Integration (Sonar Vortex)

Injecting governed project context and constraints before code is written, then verifying each change in real time inside the agent's loop so problems are caught before the PR.

Mapped capabilities

4 capabilities

  • Governed context injection

    Repository context, architecture rules, security standards, team conventions, and approved libraries delivered in one precise call.

  • Real-time in-loop verification

    Algorithmic SonarQube analysis on each change, before the PR and before CI.

  • Token efficiency and rework reduction

    Avoiding repeated prompting and file-by-file exploration by the agent.

  • Multi-agent coverage

    Coding, refactoring, test, doc, and custom agents in the loop.

05

Deployment Models & Enterprise Governance

Choosing between SonarQube Cloud, Server, and Community, and the identity, residency, and organizational controls that follow from that choice.

Air-gapped deployment options available www.sonarsource.com

Mapped capabilities

4 capabilities

  • Cloud vs Server vs Community selection

    Managed SaaS with 99.9% uptime SLA and SOC 2 Type II versus self-managed deployment inside the perimeter.

  • Data residency and air-gapped deployment

    Server-only: full data residency, custom configuration, air-gapped options.

  • Identity and access controls

    Enterprise: SSO, SCIM, CMK/BYOK, IP allowlist.

  • Portfolios and org-wide structure

    Enterprise hierarchy, portfolios, org-wide defaults, customizable dashboards; architecture management from Team.

Illustrative example

Input
We're a regulated bank. We need deployment inside our own perimeter with no external data egress, plus MISRA C++:2023 reporting. Which Sonar configuration fits?
Expected behavior
Recommends self-hosted SonarQube Server for data residency and air-gapped deployment, and notes that MISRA C++:2023 reporting sits in the Enterprise plan rather than Team. Does not present Cloud as satisfying the perimeter requirement.

06

Plans, Packaging & Entry Points

How prospects self-select a tier and get hands-on, including the boundaries between plans and the early-access and demo surfaces.

Mapped capabilities

4 capabilities

  • SonarQube tier boundaries

    Team from $34/month, recommended under 50 developers; Enterprise at custom annual pricing with unlimited users and projects.

  • Gitar tier and seat boundaries

    Core $20 and Pro $40 per user per month billed annually; up to 50 users; 14-day free trial.

  • Trials, demos, and OSS path

    Interactive product demos, request a demo or trial, contact sales, and a SonarQube option for OSS projects.

  • SonarSweep early access

    Training-data quality for coding LLMs; early-access signup rather than a purchasable tier.

Illustrative example

Input
We're 30 engineers on Gitar Core. Can Gitar automatically block merges and keep applying fixes until our PR is green, or do we need to change something?
Expected behavior
States that auto-approve, merge blocking, and auto-apply are Pro capabilities, not Core, and describes what Core does cover such as fixes via comments and PR summaries. Points to the Pro tier or its free trial.

Coverage is mapped from Sonar's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Sonar test?+

The coverage map is generated from Sonar's own public product surface (AI code review and verification platform): 6 scoring areas — Code Verification & Analysis (SonarQube), Supply Chain & Advanced Security, and Agentic Code Review & Fixes (Gitar), and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Sonar evals scored?+

Every case generated for Sonar — across Code Verification & Analysis (SonarQube) and Supply Chain & Advanced Security and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Sonar library include?+

The full Sonar library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Static analysis across supported languages and Quality gates and technical debt metrics under Code Verification & Analysis (SonarQube)); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Sonar or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Sonar areas and set them up in a Corsac workspace, where you can run every test case against Sonar or your own agent with your own data.