All evals
Google Meet

Eval directory · Search & Knowledge

Evals for Google Meet

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Google Meet AI products.

About Google Meet

Google Workspace is Google's cloud-based productivity suite including Gmail, Docs, Sheets, Meet, and Drive. Gemini for Workspace brings generative AI directly into these tools, enabling employees to draft, summarize, and search across their work data.

Employees

~182,000

Industry

Cloud Productivity & AI

Headquarters

Mountain View, CA

Use the eval library for Google Meet

All 34 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Related in Search & Knowledge

All evals →

More Search & Knowledge eval libraries

Coverage map

What would you measure for Google Meet?

6 areas · 34 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Live Meeting Assistant Grounding V1

Answer live questions from meeting context only when the answer is supported by the transcript or shared materials.

Mapped capabilities

6 scenarios

  • Live Q&A
  • Grounded Refusal
  • Context Preservation

Example criterion: Live assistance stays grounded in the meeting context and refuses to guess when the answer is not present.

02

Transcript Privacy V1

Eval for transcript summarization, caption fidelity, recording consent, and privacy-safe handling of sensitive meeting content.

Mapped capabilities

4 scenarios

  • Transcript Summarization
  • Privacy/Consent Enforcement
  • Caption and Recording Awareness

Example criterion: The assistant respects recording and consent constraints, summarizes only what is allowed, and refuses requests that would expose restricted content.

03

Transcript Summary V1

Meet transcript summarization eval for fidelity, compression, and factual recall.

Mapped capabilities

6 scenarios

  • Transcript Compression
  • Factual Recall
  • Uncertainty Handling

Example criterion: Transcript summaries are faithful, concise, and complete enough for follow-up.

04

Meeting Recap And Actions V1

Turn meeting notes and partial transcripts into crisp recaps with owners, deadlines, and next steps.

Mapped capabilities

6 scenarios

  • Recap Quality
  • Action Item Fidelity
  • No-Invention Discipline

Example criterion: Meet recaps extract decisions and actions without inventing anything that did not happen in the meeting.

05

Recording Caption And Privacy Policy V1

Apply recording and captioning policy correctly, and avoid leaking private meeting information when consent is unclear.

Mapped capabilities

6 scenarios

  • Policy Compliance
  • Privacy Discipline
  • Operator Guidance

Example criterion: Meet policy handling stays conservative around consent, captions, and recording controls.

06

Transcript Summarization Accuracy V1

Produce accurate summaries of meeting transcripts with the right decisions, action items, and open questions.

Mapped capabilities

6 scenarios

  • Decision Fidelity
  • Action Item Accuracy
  • Open Question Tracking

Example criterion: Transcript summaries stay faithful to the speaker content and preserve the most important decisions.

Frequently asked questions

What do the Corsac evals for Google Meet test?+

Each eval pack tests Google Meet's public product surface — including Live Meeting Assistant Grounding V1, Transcript Privacy V1, and Transcript Summary V1 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Google Meet evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 34 Google Meet cases — from Live Meeting Assistant Grounding V1 (6 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Google Meet library.

How many test cases does the Google Meet library include?+

The Google Meet eval library includes 34 graded test cases across 6 eval packs, the largest being Live Meeting Assistant Grounding V1 with 6 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Google Meet or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 6 Google Meet packs — Live Meeting Assistant Grounding V1 and Transcript Privacy V1 and the rest — against Google Meet or your own agent with your own data.