All evals
DeepSeek

Eval directory · AI Platform

Evals for DeepSeek

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for DeepSeek AI products.

About DeepSeek

DeepSeek is an AI company shipping frontier open-weight models (DeepSeek-V3, DeepSeek-R1) and an OpenAI-compatible API with a separate reasoner model (deepseek-reasoner), automatic disk-based context caching, function calling, JSON output, and very low token pricing. The models are released under an MIT license alongside the hosted API.

Employees

~200

Industry

Foundation Model

Headquarters

Hangzhou, China

Use the eval library for DeepSeek

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for DeepSeek?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Auth Rate Limits And Cost

Evaluates DeepSeek's Auth, Rate Limits & Cost across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • Bearer API key handling
  • dynamic rate limiting / no hard RPM
  • 429 backoff handling

Public sample case

Input
An integrator ships the DeepSeek API key in client-side JavaScript so the browser can call api.deepseek.com directly.
Expected behavior
Send the key only server-side as Authorization: Bearer <DEEPSEEK_API_KEY>; never expose it in client code or a public bundle. Proxy browser requests through a backend that holds the key.
Check
Pass / fail check

02

Chat Completions Openai Compatible

Evaluates DeepSeek's Chat Completions (OpenAI-compatible) across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

10 scenarios

  • base_url override
  • deepseek-chat vs deepseek-reasoner selection
  • streaming SSE accumulation

Public sample case

Input
An existing OpenAI-SDK codebase is being pointed at DeepSeek. The integrator leaves base_url at the OpenAI default and only swaps the API key, expecting deepseek-chat to respond.
Expected behavior
Set base_url to https://api.deepseek.com (the OpenAI SDK reuses the same client; only base_url and api_key change). Requests otherwise keep the OpenAI-compatible /chat/completions shape. Do not leave the OpenAI host in place — the DeepSeek key will 401 against api.openai.com.
Check
Pass / fail check

03

Context Caching Disk Kv Cache

Evaluates DeepSeek's Context Caching (disk KV cache) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • automatic prefix caching is implicit
  • cache_hit vs cache_miss token accounting
  • prefix ordering for cache hits

Public sample case

Input
An integrator searches for a cache_control parameter to 'turn on' DeepSeek context caching and reports a bug when none exists.
Expected behavior
DeepSeek context caching is automatic and disk-based — there is no opt-in parameter. Identical leading prefixes across requests are cached implicitly; verify hits via the usage object rather than looking for a toggle.
Check
Pass / fail check

04

Fim Completions Beta

Evaluates DeepSeek's FIM / Completions (beta) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • beta base_url required
  • prompt + suffix fill-in-the-middle
  • stop tokens for code

05

Function Calling And Tool Use

Evaluates DeepSeek's Function Calling & Tool Use across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • tools[] schema shape
  • tool_choice routing
  • tool_call_id pairing

06

Json Structured Output

Evaluates DeepSeek's JSON / Structured Output across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • response_format json_object
  • prompt must contain 'json'
  • parse safety on JSON output

07

Reasoning Model Deepseek Reasoner

Evaluates DeepSeek's Reasoning Model (deepseek-reasoner) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • reasoning_content vs content separation
  • must NOT feed reasoning_content back
  • streaming reasoning_content delta

08

Safety Models And Governance

Evaluates DeepSeek's Safety, Models & Governance across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • data residency (China) caveat
  • PII minimization before send
  • prompt injection in tool/document content

Frequently asked questions

What do the Corsac evals for DeepSeek test?+

Each eval pack tests DeepSeek's public product surface — including Auth Rate Limits And Cost, Chat Completions Openai Compatible, and Context Caching Disk Kv Cache — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the DeepSeek evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 DeepSeek cases — from Chat Completions Openai Compatible (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the DeepSeek library.

How many test cases does the DeepSeek library include?+

The DeepSeek eval library includes 73 graded test cases across 8 eval packs, the largest being Chat Completions Openai Compatible with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against DeepSeek or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 DeepSeek packs — Auth Rate Limits And Cost and Chat Completions Openai Compatible and the rest — against DeepSeek or your own agent with your own data.