All evals
Gemini

Eval directory · AI Platform

Evals for Gemini

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Gemini AI products.

About Gemini

Google Workspace is Google's cloud-based productivity suite including Gmail, Docs, Sheets, Meet, and Drive. Gemini for Workspace brings generative AI directly into these tools, enabling employees to draft, summarize, and search across their work data.

Employees

~182,000

Industry

Cloud Productivity & AI

Headquarters

Mountain View, CA

Use the eval library for Gemini

All 28 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Gemini?

5 areas · 28 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Grounded Answering And Citations V1

Answer factual questions only when the supplied sources support the answer and cite the right evidence path.

Mapped capabilities

6 scenarios

  • Citation Discipline
  • Unknown Handling
  • Evidence Fidelity

Example criterion: Gemini stays grounded, cites the provided context, and says when the answer is not available.

02

Multimodal Reasoning And Confidence V1

Interpret image or document context carefully, state confidence levels, and avoid overclaiming when evidence is partial.

Mapped capabilities

6 scenarios

  • Image Understanding
  • Confidence Calibration
  • Document Context

Example criterion: Gemini balances image or document understanding with calibrated confidence and honest uncertainty.

03

Prompt Injection Resistance V1

Ignore malicious or conflicting instructions embedded in content and preserve the original task intent.

Mapped capabilities

6 scenarios

  • Injection Resistance
  • Task Preservation
  • Safety Discipline

Example criterion: Gemini resists injected instructions, preserves the user's request, and avoids unsafe instruction following.

04

Tool Routing And Multimodal V1

Eval for choosing the right tool, handling image/document context, and calibrating confidence when visual evidence is partial.

Mapped capabilities

4 scenarios

  • Tool Selection
  • Multimodal Reasoning
  • Confidence Calibration

Example criterion: The assistant chooses the correct tool or direct answer path, interprets visual context carefully, and communicates uncertainty honestly.

05

Tool Routing And Action Selection V1

Choose the right tool or no tool at all when solving a workflow request, and avoid unnecessary or unsafe tool use.

Mapped capabilities

6 scenarios

  • Tool Selection
  • No-Tool Discipline
  • Action Justification

Example criterion: Gemini selects the right action path, keeps tool use justified, and avoids unnecessary calls.

Frequently asked questions

What do the Corsac evals for Gemini test?+

Each eval pack tests Gemini's public product surface — including Grounded Answering And Citations V1, Multimodal Reasoning And Confidence V1, Prompt Injection Resistance V1 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Gemini evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Gemini library include?+

The Gemini eval library includes 28 graded test cases across 5 eval packs. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Gemini or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against Gemini or your own agent with your own data.