All evals
Chainguard

Eval directory · Security Operations

Evals for Chainguard

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Chainguard AI products.

About Chainguard

Chainguard is a software supply chain security company that provides hardened, minimal container images with verifiable provenance. Its images and policy tooling help enterprises eliminate CVEs and meet SLSA compliance requirements in production environments.

Employees

~250

Industry

Supply Chain Security

Headquarters

Kirkland, WA

Use the eval library for Chainguard

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Chainguard?

6 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Advisory Image Operations

Evaluates Chainguard's Advisory & Image Operations across 7 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Security infrastructure eval coverage.

Mapped capabilities

7 scenarios

  • advisories list
  • no latest-dev prod
  • nightly cadence

Public sample case

Input
12 nginx advisories.
Expected behavior
Agent list+rebuild plan; no hour SLA.
Check
Pass / fail check

02

Custom Assembly Builds

Evaluates Chainguard's Custom Assembly & Builds across 6 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Security infrastructure eval coverage.

Mapped capabilities

6 scenarios

  • build apply
  • build logs
  • reject curl

Public sample case

Input
CA bundle via assembly.
Expected behavior
Agent apply with build id logged.
Check
Pass / fail check

03

Policy Gates Image Lifecycle

Evaluates Chainguard's Policy Gates & Image Lifecycle across 15 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Security infrastructure eval coverage.

Mapped capabilities

15 scenarios

  • DRY_RUN then ENFORCE
  • policy-gate check CI
  • no-eol param

Public sample case

Input
SRE runs chainctl policy-gate enable --policy=no-eol --mode=DRY_RUN --param=days=7.
Expected behavior
Agent reviews violations, ENFORCE after ticket; documents param.
Check
Pass / fail check

04

Registry Auth Pull Identity

Evaluates Chainguard's Registry Auth & Pull Identity across 15 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Security infrastructure eval coverage.

Mapped capabilities

15 scenarios

  • OIDC credential helper refresh
  • Pull token TTL for CI
  • GitHub Actions OIDC federation

05

Sbom Provenance Artifacts

Evaluates Chainguard's SBOM & Provenance Artifacts across 14 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Security infrastructure eval coverage.

Mapped capabilities

14 scenarios

  • SBOM export gap
  • SPDX vs CycloneDX
  • SBOM diff upgrade

06

Signature Attestation Verification

Evaluates Chainguard's Signature & Attestation Verification across 16 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Security infrastructure eval coverage.

Mapped capabilities

16 scenarios

  • cosign verify before pull
  • Key rotation
  • crane manifest multi-arch

Frequently asked questions

What do the Corsac evals for Chainguard test?+

Each eval pack tests Chainguard's public product surface — including Advisory Image Operations, Custom Assembly Builds, and Policy Gates Image Lifecycle — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Chainguard evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Chainguard cases — from Signature Attestation Verification (16 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Chainguard library.

How many test cases does the Chainguard library include?+

The Chainguard eval library includes 73 graded test cases across 6 eval packs, the largest being Signature Attestation Verification with 16 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Chainguard or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 6 Chainguard packs — Advisory Image Operations and Custom Assembly Builds and the rest — against Chainguard or your own agent with your own data.