All evals
Glean

Eval directory · Search & Knowledge

Evals for Glean

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Glean AI products.

About Glean

Glean is an enterprise AI work assistant that searches across every app, surface, and team to find the information employees need instantly. It builds a knowledge graph of each organization to deliver personalized, context-aware answers at work.

Employees

~500

Industry

Enterprise Search & AI

Headquarters

Palo Alto, CA

Website

glean.com

Use the eval library for Glean

All 67 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Related in Search & Knowledge

All evals →

More Search & Knowledge eval libraries

Coverage map

What would you measure for Glean?

6 areas · 67 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Assistant Agents Mcp Actions

Evaluates Glean's Assistant, Agents, MCP & Actions across 13 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Enterprise search & work AI eval coverage.

Mapped capabilities

13 scenarios

  • MCP search ACL
  • read_document denial
  • Static OAuth for enterprise ChatGPT

02

Connector Permissions Retrieval Modes

Evaluates Glean's Connector Permissions & Retrieval Modes across 13 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Enterprise search & work AI eval coverage.

Mapped capabilities

13 scenarios

  • ACL mirroring on indexed connectors
  • Live Mode per-user OAuth (Notion)
  • Indexed Notion workspace integration caveat

03

Corpus Safety Prompt Injection

Evaluates Glean's Corpus Safety & Prompt Injection across 12 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Enterprise search & work AI eval coverage.

Mapped capabilities

12 scenarios

  • Prompt injection in Confluence body
  • Injection via indexed PDF
  • Tool-use instruction in snippet

04

Cross Source Synthesis Citations

Evaluates Glean's Cross-Source Synthesis & Citations across 11 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Enterprise search & work AI eval coverage.

Mapped capabilities

11 scenarios

  • Citation faithfulness
  • Cross-source conflict surfacing
  • Multi-source synthesis without source confusion

05

Freshness Staleness Query Rewrite

Evaluates Glean's Freshness, Staleness & Query Rewrite across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Enterprise search & work AI eval coverage.

Mapped capabilities

10 scenarios

  • Ambiguous query clarification
  • Prefer newer indexed revision
  • Live fetch for freshness

06

Governance Pii Retention

Evaluates Glean's Governance, PII & Retention across 8 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Enterprise search & work AI eval coverage.

Mapped capabilities

8 scenarios

  • SSN redaction in snippet
  • Minimize home address exposure
  • GDPR access scope

Frequently asked questions

What do the Corsac evals for Glean test?+

Each eval pack tests Glean's public product surface — including Assistant Agents Mcp Actions, Connector Permissions Retrieval Modes, Corpus Safety Prompt Injection — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Glean evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Glean library include?+

The Glean eval library includes 67 graded test cases across 6 eval packs. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Glean or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against Glean or your own agent with your own data.