All evals
Antithesis

Eval directory · AI Platform

Evals for Antithesis

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Antithesis AI products.

About Antithesis

Antithesis is an autonomous software-testing platform that deterministically simulates complete systems, injects faults, and produces reproducible bug reports.

Industry

Software Testing

Headquarters

Vienna, VA

Use the eval library for Antithesis

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Antithesis?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Autonomous Exploration

Evaluates Antithesis' Autonomous State-Space Exploration across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Deterministic Testing eval coverage.

Mapped capabilities

9 scenarios

  • exploration is autonomous, not scripted scenarios
  • branching shares prefixes (multiverse search)
  • coverage / reachability feedback guides search

Public sample case

Input
A QA lead asks which exact failure scenarios to enumerate for Antithesis to run.
Expected behavior
Provide the system, a workload that exposes choices, and properties — then let the platform autonomously search the state space for property violations, rather than hand-enumerating scenarios. The platform's differentiator is finding bugs the team did not think to script. Enumerated scenarios are a…
Check
Pass / fail check

02

Cicd Auth And Governance

Evaluates Antithesis' CI/CD, Auth & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Deterministic Testing eval coverage.

Mapped capabilities

10 scenarios

  • gate releases on a CI Antithesis run
  • CI budget vs nightly soak split
  • API token scoping and storage

Public sample case

Input
Operator wants every release candidate to run an Antithesis test before promotion to production.
Expected behavior
Wire an Antithesis run into the CI/CD pipeline as a gate: build the pinned SUT image, trigger a run with a defined budget, and block promotion on any property violation or on a Sometimes guard never firing (coverage regression). [REQUIRES-VERIFICATION] for the exact CI trigger/integration mechanism…
Check
Pass / fail check

03

Deterministic Simulation And Reproducibility

Evaluates Antithesis' Deterministic Simulation & Reproducibility across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Deterministic Testing eval coverage.

Mapped capabilities

9 scenarios

  • all nondeterminism flows through the hypervisor
  • bug reproduces from the recorded seed
  • time-travel to the moment before failure

Public sample case

Input
Operator's service-under-test reads the wall clock via the OS, opens a TCP socket, and seeds its RNG from /dev/urandom. They run it under Antithesis and expect a found bug to replay identically.
Expected behavior
Run the entire system inside the Antithesis deterministic hypervisor so the clock, network, thread scheduling, and randomness are all controlled by the simulator — that is what makes a run reproducible. Do NOT bypass the simulated environment (e.g., calling out to a real external clock/API), becaus…
Check
Pass / fail check

04

Fault Injection

Evaluates Antithesis' Fault Injection across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Deterministic Testing eval coverage.

Mapped capabilities

9 scenarios

  • network partition between nodes
  • process crash and restart
  • network latency and message reordering

05

Sdk Assertions

Evaluates Antithesis' SDK Assertions (Always / Sometimes / Reachable) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Deterministic Testing eval coverage.

Mapped capabilities

9 scenarios

  • Always assertion encodes a safety invariant
  • Sometimes assertion guards against vacuity
  • Reachable / unreachable distinctions

06

Sut Setup Containers

Evaluates Antithesis' System-Under-Test Setup (Containers) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Deterministic Testing eval coverage.

Mapped capabilities

9 scenarios

  • Docker Compose defines the SUT topology
  • container images pinned by digest
  • readiness / health signaling to the simulator

07

Test Composer And Workloads

Evaluates Antithesis' Test Composer & Workload Drivers across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Deterministic Testing eval coverage.

Mapped capabilities

9 scenarios

  • workload as a driver, not a fixed script
  • parallel / composable test commands
  • workload randomness uses platform-provided entropy

08

Triage Reports And Debugging

Evaluates Antithesis' Triage Reports & Multiverse Debugging across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Deterministic Testing eval coverage.

Mapped capabilities

9 scenarios

  • report ties failure to root cause
  • minimized reproduction over raw trace
  • compare passing vs failing branches

Frequently asked questions

What do the Corsac evals for Antithesis test?+

Each eval pack tests Antithesis's public product surface — including Autonomous Exploration, Cicd Auth And Governance, and Deterministic Simulation And Reproducibility — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Antithesis evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Antithesis cases — from Cicd Auth And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Antithesis library.

How many test cases does the Antithesis library include?+

The Antithesis eval library includes 73 graded test cases across 8 eval packs, the largest being Cicd Auth And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Antithesis or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Antithesis packs — Autonomous Exploration and Cicd Auth And Governance and the rest — against Antithesis or your own agent with your own data.