All evals
Google Workspace

Eval directory · Document Agents

Evals for Google Workspace

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Google Workspace AI products.

About Google Workspace

Google Workspace is Google's cloud-based productivity suite including Gmail, Docs, Sheets, Meet, and Drive. Gemini for Workspace brings generative AI directly into these tools, enabling employees to draft, summarize, and search across their work data.

Employees

~182,000

Industry

Cloud Productivity & AI

Headquarters

Mountain View, CA

Use the eval library for Google Workspace

All 32 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Related in Document Agents

All evals →

More Document Agents eval libraries

Coverage map

What would you measure for Google Workspace?

6 areas · 32 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Calendar Scheduling Conflicts V1

Schedule meetings across time zones, detect conflicts, and ask for clarification when constraints cannot be satisfied.

Mapped capabilities

6 scenarios

  • Conflict Detection
  • Time Zone Accuracy
  • Clarifying Questions

Example criterion: Scheduling decisions respect time zones, attendee availability, and explicit constraints.

02

Docs To Action V1

Extract action items, owners, and next steps from Workspace documents without inventing missing details.

Mapped capabilities

6 scenarios

  • Action Extraction
  • Missing Context Discipline
  • Sensitive Doc Handling

Example criterion: Workspace doc responses stay grounded, preserve ownership, and avoid fabricating follow-up work.

03

Drive Search Share Governance V1

Find the right Drive artifact, respect permission boundaries, and handle external-sharing requests with discipline.

Mapped capabilities

6 scenarios

  • Document Discovery
  • Permission Discipline
  • External Sharing Safety

Example criterion: Drive responses find the right file, avoid bad shares, and keep permission changes auditable.

04

Drive Sharing Governance V1

Eval for Drive search, file selection, permission discipline, and sharing governance under realistic collaboration requests.

Mapped capabilities

4 scenarios

  • Source of Truth Selection
  • Permission Hygiene
  • Version Drift Detection

Example criterion: The assistant finds the right file, respects access boundaries, and escalates when a request would widen permissions unsafely.

05

Gmail Triage And Reply V1

Prioritize inbox messages, draft appropriate replies, and escalate only high-risk mail with strong justification.

Mapped capabilities

6 scenarios

  • Inbox Triage
  • Reply Drafting
  • Threat Detection

Example criterion: Inbox responses are concise, correctly classified, and careful about phishing, legal threats, and customer urgency.

06

Gmail Calendar Coordination V1

Eval for Gmail triage, response drafting, and calendar scheduling when multiple stakeholders and time zones are involved.

Mapped capabilities

4 scenarios

  • Inbox Triage
  • Calendar Coordination
  • Response Drafting

Example criterion: The assistant prioritizes urgent mail correctly, drafts usable replies, and schedules meetings without double-booking or timezone mistakes.

Frequently asked questions

What do the Corsac evals for Google Workspace test?+

Each eval pack tests Google Workspace's public product surface — including Calendar Scheduling Conflicts V1, Docs To Action V1, Drive Search Share Governance V1 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Google Workspace evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Google Workspace library include?+

The Google Workspace eval library includes 32 graded test cases across 6 eval packs. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Google Workspace or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against Google Workspace or your own agent with your own data.