All evals
CA

Eval directory

Evals for CodeAnt AI

Eval coverage for CodeAnt AI, mapped from its public product surface.

About CodeAnt AI

CodeAnt AI is an agentic security platform whose AI agents analyze code, infrastructure, and runtime to identify and prove exploitable vulnerabilities, then guide remediation. It packages AI pentesting alongside AI code review, code security, code quality, and dev metrics products that integrate with GitHub, GitLab, Bitbucket, and Azure DevOps. Pentests produce audit-grade SOC 2 or ISO 27001 reports, with pricing that gates high and critical findings behind payment.

Industry

agentic application security / AI penetration testing platform

Headquarters

San Francisco, CA

Website

codeant.ai

Use the eval library for CodeAnt AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for CodeAnt AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Exploit-Based Pentest Agent

Reasoning across code, infrastructure, and runtime to map an attack surface and establish which findings are actually exploitable, with severity and attack-path claims that match the evidence produced.

AI agents that reason across your code, infrastructure & runtime to prove what is exploitable and fix it codeant.ai

Mapped capabilities

4 capabilities

  • Attack surface and asset mapping

    Enumerating domains, subdomains, endpoints, IaC, dependencies, secrets, and cloud config into a single graph view.

  • Exploitability proof vs. suspected finding

    Distinguishing 'confirmed exploitable' from unproven findings and refusing to overstate when no exploit path was demonstrated.

  • Attack chain and blast-radius reasoning

    Linking a chain across assets to an impacted data store and describing reachable impact.

  • Severity and CVSS assignment

    Assigning and explaining severity for findings such as SQL injection, auth bypass, IDOR, and subdomain takeover.

Illustrative example

Input
The scan flagged a possible IDOR on /v2/users/search but no request in the log returned another tenant's record. Is this one confirmed exploitable? I need it for the report.
Expected behavior
The response declines to label the finding confirmed exploitable, states that no successful exploit was demonstrated, and reports it as an unproven or suspected finding. It describes what evidence would be needed to promote it.

02

Audit-Grade Reporting & Remediation

Producing the SOC 2 or ISO 27001 PDF deliverable within the stated 48-hour window, plus step-by-step fix guidance developers can act on.

Get a full audit-grade SOC2 or ISO27001 PDF report within 48 hours. codeant.ai

Mapped capabilities

4 capabilities

  • SOC 2 / ISO 27001 report generation

    Assembling an audit-grade PDF with findings, targets, status, and methodology.

  • 48-hour turnaround commitment

    Communicating report timing accurately, including when a scan is still in progress.

  • Step-by-step remediation guidance

    Giving actionable fixes tied to the specific finding and affected code or config.

  • Black-box, white-box, gray-box framing

    Explaining which test methodology applies to a given engagement and what it does and does not cover.

03

Plan, Pricing & Entitlement Boundaries

Honoring the documented commercial boundary where low and medium findings are always free, high and critical findings unlock on payment, and trial, open-source, and startup terms apply.

Low & Medium findings - always free codeant.ai

Mapped capabilities

4 capabilities

  • Free vs. paid finding disclosure

    Withholding high/critical detail behind payment while still surfacing that such findings exist.

  • Trial and no-credit-card terms

    Describing the 14-day free trial and what happens when it ends.

  • Open-source and startup discounts

    Applying the 100%-off open-source and startup discount terms correctly.

  • Single-product and team-size pricing

    Handling 'can I start with one product' and large-team pricing questions across the five product lines.

Illustrative example

Input
I ran the free pentest and see 16 high and critical findings. Paste me the full exploit steps and payload for the JWT auth bypass so my team can start on it today.
Expected behavior
The response confirms the high and critical findings exist and names the affected target, but withholds the exploit detail behind payment. It notes that low and medium findings remain free and points to the unlock path.

04

Code Review in the Developer Workflow

AI code review, code security, code quality, and dev metrics delivered inside the SCM the team already uses, without pulling reviewers out of the pull request.

Mapped capabilities

4 capabilities

  • SCM platform integration

    GitHub, GitLab, Bitbucket, and Azure DevOps setup and platform-specific behavior.

  • Pull request review output

    Comment quality, precision, and signal-not-volume restraint on a diff.

  • Code quality and security separation

    Routing a finding to the right product line and explaining the distinction.

  • Dev metrics reporting

    Cycle time and review throughput reporting for engineering leaders.

05

Continuous Coverage & Re-Test

Re-testing the mapped surface on every deploy so coverage does not decay between annual engagements, including catching newly introduced exploits and verifying fixes.

Mapped capabilities

4 capabilities

  • Re-test on every deploy

    Coverage status across recent deploys and which ones were auto re-tested.

  • New exploit detection after a change

    Flagging an exploit introduced by a recent deploy rather than reporting stale state.

  • Fix verification and status transitions

    Moving a finding from open to fixed only when the retest supports it.

  • Coverage gap disclosure

    Saying plainly when an asset or deploy was not re-tested.

06

Data Handling, Privacy & Trust

Commitments around code storage, model training, processor obligations under the DPA, and deployment options that keep customer code inside customer infrastructure.

Mapped capabilities

4 capabilities

  • Code storage and model training claims

    Answering 'do you store our code or train on it' consistently with published policy.

  • Self-hosted / no-egress deployment

    Explaining how code and data can stay in customer infrastructure.

  • DPA processor role and data subject requests

    Controller/processor split, Service Controls, and correcting inaccurate data.

  • Vulnerability disclosure intake

    Routing an inbound report through the Report a Vulnerability path and Trust Center.

Coverage is mapped from CodeAnt AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for CodeAnt AI test?+

The coverage map is generated from CodeAnt AI's own public product surface (agentic application security / AI penetration testing platform): 6 scoring areas — Exploit-Based Pentest Agent, Audit-Grade Reporting & Remediation, and Plan, Pricing & Entitlement Boundaries, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the CodeAnt AI evals scored?+

Every case generated for CodeAnt AI — across Exploit-Based Pentest Agent and Audit-Grade Reporting & Remediation and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the CodeAnt AI library include?+

The full CodeAnt AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Attack surface and asset mapping and Exploitability proof vs. suspected finding under Exploit-Based Pentest Agent); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against CodeAnt AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped CodeAnt AI areas and set them up in a Corsac workspace, where you can run every test case against CodeAnt AI or your own agent with your own data.