All evals
Mistral

Eval directory

Evals for Mistral

Eval coverage for Mistral, mapped from its public product surface.

About Mistral

Mistral AI is a European foundation-model company offering open-weight and commercial models (Mistral Large, Codestral, Pixtral) via La Plateforme, plus Le Chat, embeddings, fine-tuning, and agents — with a strong emphasis on EU data residency.

Employees

~250

Industry

Foundation Model

Headquarters

Paris, France

Website

mistral.ai

Use the eval library for Mistral

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Mistral?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Vibe: long-horizon assistant work

Vibe as an AI agent for multi-step knowledge work — briefing, research, synthesis, and scheduling — grounded in connected tools and persistent memory.

Mapped capabilities

4 capabilities

  • Enterprise knowledge search and grounding

    Answers drawn from connected sources, with attribution and explicit gaps rather than invented detail.

  • Document and report synthesis

    Turning source material into finished drafts, briefs, and frameworks in one conversation.

  • Multi-step task scheduling

    Decomposing a long-horizon request into ordered steps and reporting progress and blockers.

  • Persistent memory and reusable skills

    Carrying context and saved skills across sessions without leaking stale or unrelated state.

02

Vibe: autonomous coding

The coding agent that writes, tests, and deploys across a codebase from the CLI, IDE, and web.

Vibe writes, tests, and deploys across your codebase, while your developers stay on architecture mistral.ai

Mapped capabilities

4 capabilities

  • Code changes across a repository

    Multi-file edits that match surrounding conventions and stay inside the requested scope.

  • Test and verification loops

    Running tests, reporting real failures, and not claiming success on unverified work.

  • Deploy and irreversible actions

    Confirming before outward-facing or hard-to-reverse steps unless already authorized.

  • Surface parity: CLI, IDE, web

    Consistent behavior and session limits across the three documented coding interfaces.

03

Model and API capability catalog

Selecting the right endpoint across text, reasoning, coding, agentic, multimodal, OCR, voice, embedding, and classifier APIs.

Mapped capabilities

4 capabilities

  • Model selection for a stated task

    Recommending a documented model id that matches the capability tag the task needs.

  • OCR and document understanding

    OCR 4 for extraction versus Document AI, and what each returns.

  • Voice and audio generation

    Voxtral TTS on /v1/audio/speech, including voice cloning framing.

  • Open, Premier, and Labs licensing

    Distinguishing license tiers and their deployment implications.

Illustrative example

Input
We need to pull text out of 12,000 scanned invoices. Which Mistral model should we use, and roughly what will that cost?
Expected behavior
Recommends OCR 4 (mistral-ocr-4-0) and prices it per 1,000 pages at $4 for OCR or $5 for Document AI, noting that the estimate depends on total page count rather than document count and that a page count is needed to finalize.

04

Pricing and billing mechanics

Plan tiers and API rate math, including discounts, alternate billing units, currency, and tax presentation.

including regional data processing controls, system-level SLAs, increased rate limits, and premium support mistral.ai

Mapped capabilities

4 capabilities

  • Consumer and team plan tiers

    Free, Pro $14.99/mo, Team $24.99/user/mo, Education $5.99, and Enterprise contact-sales.

  • Batch and cached-token discounts

    50% batch processing and 90% cached input token reductions applied correctly.

  • Non-token billing units

    Per-1000-page OCR pricing and per-1k-character audio generation.

  • Currency, tax, and region display

    USD/EUR toggle, excluding/including tax, and the supported country list.

Illustrative example

Input
I'm sending 2M input tokens to mistral-medium-latest through the Batch API, and about 90% of them are cache hits. What do I pay for input?
Expected behavior
Uses the published $1.5 per 1M input rate for Mistral Medium 3.5, applies the 50% batch discount and the 90% cached-input discount to the cached share, and states whether it assumes the two discounts stack rather than inventing an unlisted rate.

05

Studio: building and running agents

The surface for building, testing, and running AI agents and apps on top of Mistral models.

Custom models. Custom agents. Custom workflows. Audit logs. SAML SSO. White label. mistral.ai

Mapped capabilities

4 capabilities

  • Agent construction and configuration

    Assembling an agent from models, tools, and instructions.

  • Testing before rollout

    Exercising an agent against sample inputs prior to running it in production.

  • Running agents in production

    Operating deployed agents that route and resolve repeating work.

  • Connector and tool integration

    Wiring the documented 100+ connectors into an agent's reach.

06

Forge and enterprise deployment

Custom model lifecycle plus the governance and deployment controls that enterprise buyers evaluate.

Strict data isolation, controlled training pipelines, and auditable customization workflows aligned to your compliance policies. mistral.ai

Mapped capabilities

4 capabilities

  • Data preparation and synthetic coverage

    Domain examples, edge cases, and policy-bound scenarios from proprietary data.

  • Training, alignment, and KPI evaluation

    Full lifecycle through post-training RL, evaluated against enterprise KPIs rather than generic benchmarks.

  • Deployment flexibility and data isolation

    Private deployments, regional inference endpoints, and no single-cloud lock-in.

  • Governance controls

    Audit logs, SAML SSO, white label, domain verification, and data export.

Coverage is mapped from Mistral's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Mistral test?+

The coverage map is generated from Mistral's own public product surface (frontier AI models and enterprise AI platform (assistants, agents, model customization)): 6 scoring areas — Vibe: long-horizon assistant work, Vibe: autonomous coding, and Model and API capability catalog, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Mistral evals scored?+

Every case generated for Mistral — across Vibe: long-horizon assistant work and Vibe: autonomous coding and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Mistral library include?+

The full Mistral library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Enterprise knowledge search and grounding and Document and report synthesis under Vibe: long-horizon assistant work); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Mistral or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Mistral areas and set them up in a Corsac workspace, where you can run every test case against Mistral or your own agent with your own data.