All evals
Docker

Eval directory · AI Platform

Evals for Docker

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Docker AI products.

About Docker

Docker is the container platform — Docker Engine, Docker Desktop, Docker Hub registry, Docker Build Cloud for managed cloud builders, Docker Scout for image vulnerability scanning and supply-chain policy, Docker Compose for multi-container dev, and Docker Model Runner for local LLM inference. Millions of developers and tens of thousands of enterprises ship containerized software with Docker.

Employees

~600

Industry

Developer Infrastructure

Headquarters

Palo Alto, CA

Use the eval library for Docker

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Docker?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Docker Build Cloud

Evaluates Docker's Docker Build Cloud across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Container Platform eval coverage.

Mapped capabilities

9 scenarios

  • cloud driver provisioning
  • native ARM cross-build
  • shared team cache

Public sample case

Input
Operator wants Docker Build Cloud builds from CI. CI currently uses the default 'docker' driver which only builds for the runner's architecture.
Expected behavior
Run 'docker buildx create --driver cloud <org>/<builder-name> --use'. Subsequent 'docker buildx build' streams the build to the cloud builder endpoint; native ARM and amd64 stages run on dedicated hardware (no QEMU). Authenticate to Docker Hub first (docker login).
Check
Pass / fail check

02

Docker Compose And Local Dev

Evaluates Docker's Docker Compose & Local Dev across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Container Platform eval coverage.

Mapped capabilities

9 scenarios

  • depends_on with healthcheck condition
  • compose watch for dev sync
  • profiles for environment variants

Public sample case

Input
compose.yaml: web depends_on: [db]. App starts and crashes because Postgres isn't ready to accept connections yet.
Expected behavior
Use long-form depends_on with condition: service_healthy and add a healthcheck to the db service (pg_isready or psql query). web waits for healthcheck pass, not just container start. Short-form depends_on only waits for container start, not readiness.
Check
Pass / fail check

03

Docker Desktop And Extensions

Evaluates Docker's Docker Desktop & Extensions across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Container Platform eval coverage.

Mapped capabilities

9 scenarios

  • resource limits saturate dev VM
  • file sharing mode VirtioFS vs gRPC FUSE
  • Kubernetes integration

Public sample case

Input
Developer reports 'docker compose up' fails halfway with OOM. Desktop is set to 4 GB RAM, 2 CPU; the stack runs 12 services.
Expected behavior
Increase Desktop resources via Settings → Resources → Advanced (or settings.json: memoryMiB, cpus). 8-16 GB is typical for multi-service dev. Verify via 'docker info' showing the new limit. Also confirm the host has headroom — Desktop allocates from the host RAM.
Check
Pass / fail check

04

Docker Engine Containers Runtime

Evaluates Docker's Docker Engine, Containers & Runtime across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Container Platform eval coverage.

Mapped capabilities

9 scenarios

  • HEALTHCHECK directive
  • bridge vs overlay network
  • bind mount vs named volume

05

Docker Hub And Registry

Evaluates Docker's Docker Hub & Registry across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Container Platform eval coverage.

Mapped capabilities

9 scenarios

  • tag mutability vs digest pinning
  • access token vs password
  • anonymous pull rate limit

06

Docker Scout

Evaluates Docker's Docker Scout across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Container Platform eval coverage.

Mapped capabilities

9 scenarios

  • scout cves on image
  • SBOM generation
  • policy evaluation

07

Dockerfile And Image Build

Evaluates Docker's Dockerfile & Image Build across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Container Platform eval coverage.

Mapped capabilities

9 scenarios

  • multi-stage build leaks build deps
  • BuildKit secret mount vs ARG
  • cache mount for package manager

08

Model Runner And Safety Governance

Evaluates Docker's Docker Model Runner & Safety/Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Container Platform eval coverage.

Mapped capabilities

10 scenarios

  • docker model pull from ai/ namespace
  • OpenAI-compatible /engines/v1
  • quantization choice

Frequently asked questions

What do the Corsac evals for Docker test?+

Each eval pack tests Docker's public product surface — including Docker Build Cloud, Docker Compose And Local Dev, and Docker Desktop And Extensions — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Docker evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Docker cases — from Model Runner And Safety Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Docker library.

How many test cases does the Docker library include?+

The Docker eval library includes 73 graded test cases across 8 eval packs, the largest being Model Runner And Safety Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Docker or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Docker packs — Docker Build Cloud and Docker Compose And Local Dev and the rest — against Docker or your own agent with your own data.