All evals
Mistral AI

Eval directory · AI Platform

Evals for Mistral AI

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Mistral AI AI products.

About Mistral AI

Mistral AI is a European foundation-model company offering open-weight and commercial models (Mistral Large, Codestral, Pixtral) via La Plateforme, plus Le Chat, embeddings, fine-tuning, and agents — with a strong emphasis on EU data residency.

Employees

~250

Industry

Foundation Model

Headquarters

Paris, France

Website

mistral.ai

Use the eval library for Mistral AI

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Mistral AI?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Mistral Chat Completions And Streaming

Evaluates Mistral AI's Chat Completions & Streaming across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

10 scenarios

  • streaming finish_reason length
  • SSE [DONE] terminator
  • stop sequence handling

Public sample case

Input
Agent streams /v1/chat/completions with stream=true and max_tokens=256 for a long answer; the final SSE chunk reports finish_reason='length'.
Expected behavior
Detect finish_reason='length' on the terminal chunk and treat the answer as truncated — surface a partial-completion to the caller or continue by appending the assistant text and re-requesting. Never present truncated output as complete.
Check
Pass / fail check

02

Mistral Embeddings And Retrieval

Evaluates Mistral AI's Embeddings & Retrieval across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • consistent model across index
  • fixed dimensionality assumption
  • normalization for cosine

Public sample case

Input
Team builds a retrieval index with mistral-embed but later embeds new queries with a different embedding model, then compares vectors.
Expected behavior
All vectors in an index must come from the same embedding model; query vectors must be produced by mistral-embed if the index was built with mistral-embed. Mixing models makes similarity meaningless.
Check
Pass / fail check

03

Mistral Fine Tuning And Model Customization

Evaluates Mistral AI's Fine-tuning & Model Customization across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • training file format validation
  • train/validation split
  • hyperparameter overfit

Public sample case

Input
Operator uploads a JSONL training file for a fine-tuning job where ~8% of lines are malformed (missing the assistant turn).
Expected behavior
Validate the training file format (one chat per line with the required roles) before creating the job; pre-checking avoids burning a failed job. Fix or drop malformed lines and re-validate.
Check
Pass / fail check

04

Mistral Function Calling And Tool Use

Evaluates Mistral AI's Function Calling & Tool Use across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • tool_choice=any forces a call
  • tool_choice=none suppresses calls
  • parallel tool_calls id pairing

05

Mistral Json Mode And Structured Output

Evaluates Mistral AI's JSON Mode & Structured Output across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • json_object syntax-only guarantee
  • json_schema strict conformance
  • json mode requires instruction

06

Mistral Le Chat Agents And Connectors

Evaluates Mistral AI's Le Chat / Agents & Connectors across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • agent handoff context
  • web search connector grounding
  • code interpreter sandbox trust

07

Mistral Models Versioning And Deployment

Evaluates Mistral AI's Models, Versioning & Deployment across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • version pinning vs alias drift
  • open-weights vs API model choice
  • code model for code tasks

08

Mistral Safety Moderation And Governance

Evaluates Mistral AI's Safety, Moderation & Governance across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • pre-screen UGC with moderation
  • category-score vs flag gating
  • safe_prompt is not a classifier

Frequently asked questions

What do the Corsac evals for Mistral AI test?+

Each eval pack tests Mistral AI's public product surface — including Mistral Chat Completions And Streaming, Mistral Embeddings And Retrieval, and Mistral Fine Tuning And Model Customization — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Mistral AI evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Mistral AI cases — from Mistral Chat Completions And Streaming (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Mistral AI library.

How many test cases does the Mistral AI library include?+

The Mistral AI eval library includes 73 graded test cases across 8 eval packs, the largest being Mistral Chat Completions And Streaming with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Mistral AI or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Mistral AI packs — Mistral Chat Completions And Streaming and Mistral Embeddings And Retrieval and the rest — against Mistral AI or your own agent with your own data.