All evals
Anthropic

Eval directory · AI Platform

Evals for Anthropic

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Anthropic AI products.

About Anthropic

Anthropic is an AI safety company and the maker of Claude. Its API exposes the Claude model family (Opus, Sonnet, Haiku) with tool use, prompt caching, extended thinking, batch processing, vision, the Files and Memory tools, and the Claude Agent SDK.

Employees

~1,000

Industry

Foundation Model

Headquarters

San Francisco, CA

Use the eval library for Anthropic

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Anthropic?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Batch Api

Evaluates Anthropic's Batch API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • create batch with custom_id
  • 24-hour processing window
  • results retrieval

Public sample case

Input
Agent submits 5000 requests to POST /v1/messages/batches, each carrying a unique custom_id matching its row in the operator's dataset.
Expected behavior
Build requests[] with custom_id and params (model, messages, max_tokens, tools, etc.) per row. custom_id is the only way to map results back to source rows — pick a stable, unique value (e.g., row_uuid).
Check
Pass / fail check

02

Extended Thinking

Evaluates Anthropic's Extended Thinking across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • thinking config shape
  • thinking blocks before text
  • redacted_thinking preservation

Public sample case

Input
Agent enables extended thinking with thinking={type:'enabled', budget_tokens:10000}, max_tokens=16000.
Expected behavior
Both fields are required when enabling. budget_tokens reserves thinking capacity within max_tokens (text output gets max_tokens - thinking tokens used). Verify thinking + text totals stay under max_tokens.
Check
Pass / fail check

03

Files Api And Citations

Evaluates Anthropic's Files API & Citations across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • upload returns file_id
  • document source via file_id
  • citation block structure

Public sample case

Input
Operator uploads a 12 MB invoice PDF via POST /v1/files and receives a file_id.
Expected behavior
Persist file_id with its source-document mapping. Reference it in Messages calls via source={type:'file', file_id:'<id>'} on a document content block. Do not re-upload identical files (use prior file_id) — but plan for deletion lifecycle.
Check
Pass / fail check

04

Memory Tool And Context Editing

Evaluates Anthropic's Memory Tool & Context Editing across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • beta header required
  • view command lists memory
  • create writes new file

05

Messages Api And Streaming Sse

Evaluates Anthropic's Messages API & Streaming SSE across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

10 scenarios

  • SSE event ordering
  • ping events
  • stop_reason=max_tokens

06

Prompt Caching

Evaluates Anthropic's Prompt Caching across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • cache_control on system
  • cache_control on tools
  • 4-breakpoint limit

07

Refusals Safety And Agent Sdk

Evaluates Anthropic's Refusals, Safety & Agent SDK / Claude Code across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • AUP-prohibited request handling
  • jailbreak attempt resistance
  • refusal content shape

08

Tool Use And Schema Validation

Evaluates Anthropic's Tool Use & Schema Validation across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • tool_choice=auto
  • tool_choice=tool forces single tool
  • parallel tool calls

Frequently asked questions

What do the Corsac evals for Anthropic test?+

Each eval pack tests Anthropic's public product surface — including Batch Api, Extended Thinking, and Files Api And Citations — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Anthropic evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Anthropic cases — from Messages Api And Streaming Sse (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Anthropic library.

How many test cases does the Anthropic library include?+

The Anthropic eval library includes 73 graded test cases across 8 eval packs, the largest being Messages Api And Streaming Sse with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Anthropic or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Anthropic packs — Batch Api and Extended Thinking and the rest — against Anthropic or your own agent with your own data.