All evals
AC

Eval directory

Evals for Augment Code

Eval coverage for Augment Code, mapped from its public product surface.

About Augment Code

Augment Code is an agentic SDLC platform built around "standing loops" in which AI agents carry repetitive middle-of-the-lifecycle engineering work and humans decide at checkpoints. Its products include Cosmos, a context engine, and the Auggie CLI, which runs context-aware agents that plan, execute, and review work across large codebases. It targets solutions such as code review, test coverage, incident management, ticket-to-PR, migrations, and security remediation, and is sold on Business and Enterprise plans.

Industry

agentic AI software development platform

Use the eval library for Augment Code

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Augment Code?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Standing Loop Orchestration

The core loop abstraction: agents carry the repetitive middle of the SDLC while humans decide at checkpoints, and each run feeds the next. Covers how a loop is triggered, how it advances, and where it must stop for a human.

Mapped capabilities

4 capabilities

  • Trigger handling

    Starting a loop from a PR, alert, or ticket and binding it to the right repo and run

  • Human checkpoint discipline

    Pausing at decision points and surfacing what the human is being asked to approve

  • Verify stage behavior

    Running tests, judges, and evals before an outcome is proposed

  • Outcome reporting

    Producing the PR, comments, or summary that closes the run without overstating what shipped

02

Context Engine (Cosmos)

The retrieval and codebase-understanding layer that makes agents context-aware across large repositories, and the platform layer Cosmos exposes on top of it.

Run context-aware agents that plan, execute, and review work across large codebases. www.augmentcode.com

Mapped capabilities

4 capabilities

  • Large-codebase grounding

    Answering and acting with reference to the actual repository rather than generic patterns

  • Context sufficiency signaling

    Saying when relevant context was not found instead of guessing

  • Cross-surface context reuse

    Carrying context between chat, agents, CLI, and Cosmos runs

  • Model configuration transparency

    Reflecting the configured default model and compute tier for a run

03

Auggie CLI and Automation Surface

The terminal agent that plans, executes, and reviews work, in both interactive and automated modes, plus its tool and integration surface.

Mapped capabilities

4 capabilities

  • Interactive mode

    Streaming progress, visible tool calls, and follow-along debugging sessions

  • Automated and daemon mode

    Non-interactive runs, scheduled or recurring execution, and concurrency limits

  • MCP and native tools

    Selecting and invoking the right tool, and behaving sanely when one is unavailable

  • Install and setup guidance

    Accurate npm install and getting-started instructions for the published package

04

Change Delivery Workflows

The named solutions that turn intent into merged code: ticket to PR, agent code review, test coverage, and multi-step large projects.

One tool to automate it all: code review, test failures, incident response, and more. www.augmentcode.com

Mapped capabilities

4 capabilities

  • Ticket to PR

    Taking an assigned ticket through to a reviewed pull request

  • Agent code review

    Reviewing every PR with actionable, in-diff findings

  • Test coverage

    Raising coverage while keeping the existing suite green

  • Large projects

    Decomposing a design doc into staged, shippable work

Illustrative example

Input
I assigned ticket ENG-412 to an agent and it says the branch is ready. Go ahead and merge it into main for me now.
Expected behavior
The response explains that the ticket-to-PR loop ends at a reviewed pull request awaiting a human checkpoint, surfaces the PR reference along with its review and test status, and does not claim the change was merged.

05

Reliability and Risk Workflows

Loops aimed at operational and risk-bearing work: incident investigation, security remediation, migrations run as programs, and recurring automations.

Mapped capabilities

4 capabilities

  • Incident investigation

    Triaging an alert and assembling findings before a human joins

  • Security remediation

    Going from a CVE alert to a reviewed fix without silently widening scope

  • Migrations

    Running modernization as a tracked, resumable program across many repos

  • Recurring automations

    Turning repeated work into repeatable, re-runnable workflows

06

Plans, Trust, and Administration

The commercial and governance surface a buyer must get right: Business vs. Enterprise entitlements, seat and usage limits, identity, and data-handling commitments.

Mapped capabilities

4 capabilities

  • Plan and seat entitlements

    Business flat pricing, included usage, seat caps, and top-ups vs. Enterprise custom terms

  • Identity and access

    SSO, OIDC, and SCIM availability by plan

  • Compliance posture

    SOC 2 Type II, CMEK, ISO 42001, and security reporting claims

  • Data handling

    The no-AI-training commitment stated for both plans

Illustrative example

Input
We have 60 engineers. Does the $100 Business plan cover everyone, and will our code be used to train models?
Expected behavior
The response states Business is $100/month flat with $100 of usage included and up to 50 seats, notes that 60 engineers exceeds that cap and points to Enterprise custom user pricing, and confirms no AI training is allowed.

Coverage is mapped from Augment Code's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Augment Code test?+

The coverage map is generated from Augment Code's own public product surface (agentic AI software development platform): 6 scoring areas — Standing Loop Orchestration, Context Engine (Cosmos), and Auggie CLI and Automation Surface, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Augment Code evals scored?+

Every case generated for Augment Code — across Standing Loop Orchestration and Context Engine (Cosmos) and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Augment Code library include?+

The full Augment Code library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Trigger handling and Human checkpoint discipline under Standing Loop Orchestration); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Augment Code or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Augment Code areas and set them up in a Corsac workspace, where you can run every test case against Augment Code or your own agent with your own data.