All evals
Cohere

Eval directory · AI Platform

Evals for Cohere

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Cohere AI products.

About Cohere

Cohere builds enterprise foundation models and the tools around them — the Command model family, best-in-class Rerank and Embed endpoints, and grounded retrieval-augmented generation with inline citations — deployable across major clouds and private VPCs.

Employees

~400

Industry

Foundation Model

Headquarters

Toronto, Canada

Website

cohere.com

Use the eval library for Cohere

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Cohere?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Chat Api And Streaming

Evaluates Cohere's Chat API & Streaming across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

10 scenarios

  • messages array role contract
  • streaming event types
  • citation-start / citation-end in stream

Public sample case

Input
Agent builds a /v2/chat request with messages[] where it places a system instruction as a trailing message with role='user' instead of role='system', expecting it to behave like a preamble.
Expected behavior
Use the documented v2 message roles: a message with role='system' for instructions/preamble, then alternating role='user'/role='assistant' turns. Do not smuggle system instructions into a user turn — that turn is treated as user content and can leak into the conversation transcript.
Check
Pass / fail check

02

Command Models And Versioning

Evaluates Cohere's Command Models & Versioning across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • explicit model id selection
  • dated version pinning
  • context window budgeting

Public sample case

Input
Production /v2/chat calls omit the model field and rely on an account default, so behavior shifts when the default Command model changes.
Expected behavior
Always pass an explicit model id (e.g., a Command-R / Command-R+ / Command-A family id) so behavior is reproducible. Treat any default-model change as a behavioral change requiring re-validation.
Check
Pass / fail check

03

Embed

Evaluates Cohere's Embed across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • input_type for documents vs query
  • classification / clustering input_type
  • embedding dimension consistency

Public sample case

Input
Agent embeds corpus passages and user queries with the same input_type, then wonders why retrieval recall is poor.
Expected behavior
Set input_type='search_document' when embedding corpus passages for indexing and input_type='search_query' when embedding the query at search time. The asymmetric input_type is required for the retrieval embedding space to align.
Check
Pass / fail check

04

Fine Tuning And Customization

Evaluates Cohere's Fine-tuning & Customization across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • dataset format validation
  • train/validation split discipline
  • fine-tune type matches task

05

Rag And Grounded Generation

Evaluates Cohere's RAG & Grounded Generation across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • documents[] grounding shape
  • inline citation spans correctness
  • answer not in documents (abstention)

06

Rerank

Evaluates Cohere's Rerank across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • rerank request shape
  • index→document mapping
  • relevance_score is not a probability

07

Safety Deployment And Governance

Evaluates Cohere's Safety, Deployment & Governance across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • safety_mode configuration
  • prompt injection from retrieved content
  • PII handling in prompts/logs

08

Tool Use And Function Calling

Evaluates Cohere's Tool Use / Function Calling across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.

Mapped capabilities

9 scenarios

  • tools schema declaration
  • tool_calls handling
  • tool_call_id ↔ tool result pairing

Frequently asked questions

What do the Corsac evals for Cohere test?+

Each eval pack tests Cohere's public product surface — including Chat Api And Streaming, Command Models And Versioning, and Embed — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Cohere evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Cohere cases — from Chat Api And Streaming (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Cohere library.

How many test cases does the Cohere library include?+

The Cohere eval library includes 73 graded test cases across 8 eval packs, the largest being Chat Api And Streaming with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Cohere or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Cohere packs — Chat Api And Streaming and Command Models And Versioning and the rest — against Cohere or your own agent with your own data.