All evals
xAI

Eval directory · AI Platform

Evals for xAI

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for xAI AI products.

About xAI

xAI builds the Grok foundation-model family and the xAI API — OpenAI-compatible chat completions, function calling, Live Search / DeepSearch real-time web grounding, Grok Vision multimodal inputs, reasoning models with a thinking-effort budget, and Grok / Aurora image generation.

Employees

~1,000

Industry

Foundation Model

Headquarters

Palo Alto, CA

Website

x.ai

Use the eval library for xAI

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for xAI?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Auth Rate Limits And Concurrency

Evaluates xAI's Auth, Rate Limits & Concurrency across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • Bearer token header shape
  • project-scoped vs root API keys
  • 429 with Retry-After header

Public sample case

Input
Client authenticates with Authorization: Bearer <XAI_API_KEY> on every request to https://api.x.ai/v1.
Expected behavior
Pull the key from environment / secret manager — never embed in source. Set Authorization: Bearer <key> exactly once per request. Verify on startup that the key resolves (small probe call) rather than on first user request. Rotate keys via the xAI console; revoke immediately on leak.
Check
Pass / fail check

02

Chat Completions Api

Evaluates xAI's Chat Completions API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • OpenAI-compatible base URL
  • messages[] role schema
  • stream=true SSE delta chunks

Public sample case

Input
Operator points an OpenAI SDK client at the xAI API by setting base_url='https://api.x.ai/v1' and model='grok-4'. A teammate leaves base_url at OpenAI's default and the call routes there instead.
Expected behavior
xAI exposes an OpenAI-compatible /v1/chat/completions endpoint at https://api.x.ai/v1. Set base_url explicitly per-client; do not rely on env-var inheritance from OPENAI_BASE_URL. Verify a debug log of the resolved base_url at startup before allowing traffic.
Check
Pass / fail check

03

Function Calling And Tool Use

Evaluates xAI's Function Calling & Tool Use across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • tools[] schema shape
  • tool_choice=auto vs required
  • parallel tool calls

Public sample case

Input
Agent declares tools=[{type:'function', function:{name, description, parameters: <JSON Schema>}}] on /v1/chat/completions and expects Grok to decide whether to call get_weather.
Expected behavior
Pass tools[] in the OpenAI-compatible shape (type='function', nested function object with name/description/parameters). parameters MUST be a valid JSON Schema object. Distinct, action-oriented descriptions enable correct routing.
Check
Pass / fail check

04

Image Generation Grok Aurora

Evaluates xAI's Image Generation (Grok / Aurora) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • endpoint shape and model id
  • prompt safety pre-screen
  • response_format: url vs b64_json

05

Live Search And Deepsearch

Evaluates xAI's Live Search / DeepSearch across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • search_parameters mode shape
  • citation surfacing to end user
  • freshness / recency filters

06

Reasoning And Thinking

Evaluates xAI's Reasoning & Thinking across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • reasoning_effort parameter
  • reasoning latency budget
  • reasoning token cost accounting

07

Safety Policy And Governance

Evaluates xAI's Safety, Policy & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

10 scenarios

  • xAI Usage Policy refusal handling
  • jailbreak attempt resistance
  • prompt injection in tool results

08

Vision Inputs Grok Vision

Evaluates xAI's Vision Inputs (Grok Vision) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • image_url content block shape
  • base64 vs URL upload tradeoff
  • image size and resolution caps

Frequently asked questions

What do the Corsac evals for xAI test?+

Each eval pack tests xAI's public product surface — including Auth Rate Limits And Concurrency, Chat Completions Api, and Function Calling And Tool Use — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the xAI evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 xAI cases — from Safety Policy And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the xAI library.

How many test cases does the xAI library include?+

The xAI eval library includes 73 graded test cases across 8 eval packs, the largest being Safety Policy And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against xAI or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 xAI packs — Auth Rate Limits And Concurrency and Chat Completions Api and the rest — against xAI or your own agent with your own data.