All evals
OpenAI

Eval directory · AI Platform

Evals for OpenAI

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for OpenAI AI products.

About OpenAI

OpenAI builds the GPT model family and the OpenAI API — Responses and Chat Completions, function calling, Structured Outputs, embeddings, fine-tuning, the Batch API, moderation, the Realtime API, and the Agents SDK — used by developers to build AI products at scale.

Employees

~3,000

Industry

Foundation Model

Headquarters

San Francisco, CA

Website

openai.com

Use the eval library for OpenAI

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for OpenAI?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Batch Api

Evaluates OpenAI's Batch API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • JSONL input + custom_id
  • 24h completion window
  • output + error files

Public sample case

Input
Operator submits 20k requests via a JSONL file to /v1/batches, each line a request with a custom_id matching a dataset row.
Expected behavior
Each input line needs a unique custom_id, a method, a url (/v1/responses or /v1/chat/completions), and a body. custom_id is the only mapping back to source rows; pick a stable unique value.
Check
Pass / fail check

02

Embeddings And Retrieval

Evaluates OpenAI's Embeddings & Retrieval across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • dimensions reduction
  • normalization for cosine
  • batch embedding inputs

Public sample case

Input
Team uses text-embedding-3-large but sets dimensions=256 to save vector-store cost, then compares to vectors stored at full dimension.
Expected behavior
All vectors in an index must share the same model and dimensions; re-embed the whole corpus when changing dimensions. Mixing dimensions makes cosine similarity meaningless.
Check
Pass / fail check

03

Fine Tuning

Evaluates OpenAI's Fine-tuning across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • training file validation
  • train/validation split
  • hyperparameter defaults

Public sample case

Input
Operator uploads a JSONL SFT file where 8% of lines are malformed (missing assistant turn).
Expected behavior
Validate the training file format (one chat per line with the required roles) before creating the job; the API surfaces validation errors but pre-checking saves a failed job. Fix or drop malformed lines.
Check
Pass / fail check

04

Function Calling And Tool Orchestration

Evaluates OpenAI's Function Calling & Tool Orchestration across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • tool_choice=required
  • parallel tool_calls pairing
  • parallel_tool_calls=false

05

Moderation And Safety

Evaluates OpenAI's Moderation & Safety across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • pre-screen user content
  • category vs score gating
  • multimodal moderation

06

Realtime Api And Reasoning Models

Evaluates OpenAI's Realtime API & Reasoning Models across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • server VAD interruption
  • function calls in voice
  • WebRTC vs WebSocket choice

07

Responses And Chat Completions

Evaluates OpenAI's Responses & Chat Completions across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

10 scenarios

  • streaming finish_reason
  • previous_response_id continuity
  • developer vs system role

08

Structured Outputs And Json Schema

Evaluates OpenAI's Structured Outputs & JSON Schema across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • strict json_schema guarantee
  • refusal field handling
  • unsupported schema keyword

Frequently asked questions

What do the Corsac evals for OpenAI test?+

Each eval pack tests OpenAI's public product surface — including Batch Api, Embeddings And Retrieval, and Fine Tuning — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the OpenAI evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 OpenAI cases — from Responses And Chat Completions (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the OpenAI library.

How many test cases does the OpenAI library include?+

The OpenAI eval library includes 73 graded test cases across 8 eval packs, the largest being Responses And Chat Completions with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against OpenAI or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 OpenAI packs — Batch Api and Embeddings And Retrieval and the rest — against OpenAI or your own agent with your own data.