All evals
Sublime Security

Eval directory · AI Platform

Evals for Sublime Security

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Sublime Security AI products.

About Sublime Security

Sublime Security is an AI-powered cloud email-security platform for Microsoft 365 and Google Workspace. It detects, triages, hunts, and responds to email threats with programmable detections and specialized security agents.

Industry

Email Security

Use the eval library for Sublime Security

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Sublime Security?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Actions And Remediation

Evaluates Sublime Security's Actions & Remediation across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI-Powered Email Security (Detection-as-Code / MQL) eval coverage.

Mapped capabilities

9 scenarios

  • auto-quarantine vs review queue
  • post-delivery clawback
  • warning banner vs hard action

Public sample case

Input
A high-confidence malware rule fires on a delivered message. Operator must decide between auto-removing it from the mailbox and routing to an analyst review queue.
Expected behavior
Bind the action to confidence: high-confidence malicious detections auto-remediate (move/remove from the mailbox) to cut dwell time, while ambiguous detections route to a review queue for human disposition. Auto-action only where the rule's precision justifies the false-positive cost of removing re…
Check
Pass / fail check

02

Attack Surface Reduction And Hunting

Evaluates Sublime Security's Attack-Surface Reduction & Hunting across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI-Powered Email Security (Detection-as-Code / MQL) eval coverage.

Mapped capabilities

9 scenarios

  • ad-hoc MQL hunt over historical mail
  • IOC pivot from one detection
  • insights / aggregate exposure view

Public sample case

Input
After IOC disclosure (a sender domain + a subject pattern), an analyst runs an ad-hoc MQL query across stored historical messages to find prior delivery of the same campaign.
Expected behavior
Express the hunt as an MQL query over the message store, scoped to a time window, combining the IOC domain with a structural signal so the hunt is precise. Export matches for remediation rather than eyeballing. Treat the hunt query as reusable — promote a high-value one to a standing rule. [REQUIRE…
Check
Pass / fail check

03

Detonation Enrichment And Ml Signals

Evaluates Sublime Security's Detonation, Enrichment & ML Signals across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI-Powered Email Security (Detection-as-Code / MQL) eval coverage.

Mapped capabilities

9 scenarios

  • link detonation verdict use
  • file detonation & evasive payloads
  • domain / sender reputation enrichment

Public sample case

Input
Operator references a link-detonation / sandbox verdict in a rule to catch credential-harvesting pages behind benign-looking URLs.
Expected behavior
Read the documented detonation verdict field and treat it as one signal among several — a clean verdict is not proof of safety (cloaking, geofencing, time-bombed redirects evade sandboxes), and a malicious verdict is high-value. Combine with sender/auth context. Account for verdicts that arrive asy…
Check
Pass / fail check

04

Feeds And Open Rule Ecosystem

Evaluates Sublime Security's Feeds & Open Rule Ecosystem across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI-Powered Email Security (Detection-as-Code / MQL) eval coverage.

Mapped capabilities

9 scenarios

  • import from open sublime-rules feed
  • feed update / sync cadence
  • local fork vs upstream divergence

05

Message Ingestion And Eml Analysis

Evaluates Sublime Security's Message Ingestion & EML Analysis across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI-Powered Email Security (Detection-as-Code / MQL) eval coverage.

Mapped capabilities

9 scenarios

  • submit raw .eml for analysis
  • MIME multipart & nested parts
  • encoding & display normalization

06

Mql Detection Authoring

Evaluates Sublime Security's MQL Detection Authoring across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI-Powered Email Security (Detection-as-Code / MQL) eval coverage.

Mapped capabilities

9 scenarios

  • sender vs reply-to mismatch rule
  • MQL type-correct attribute access
  • attachment-scoped MQL predicate

07

Platform Api Auth And Deployment

Evaluates Sublime Security's Platform, API, Auth & Deployment across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI-Powered Email Security (Detection-as-Code / MQL) eval coverage.

Mapped capabilities

10 scenarios

  • API key scoping & least privilege
  • API pagination & rate limits
  • RBAC for analysts vs admins

08

Rule Lifecycle And Tuning

Evaluates Sublime Security's Rule Lifecycle & Tuning across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI-Powered Email Security (Detection-as-Code / MQL) eval coverage.

Mapped capabilities

9 scenarios

  • backtest before enabling
  • false-positive exclusion discipline
  • passive vs active rule state

Frequently asked questions

What do the Corsac evals for Sublime Security test?+

Each eval pack tests Sublime Security's public product surface — including Actions And Remediation, Attack Surface Reduction And Hunting, and Detonation Enrichment And Ml Signals — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Sublime Security evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Sublime Security cases — from Platform Api Auth And Deployment (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Sublime Security library.

How many test cases does the Sublime Security library include?+

The Sublime Security eval library includes 73 graded test cases across 8 eval packs, the largest being Platform Api Auth And Deployment with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Sublime Security or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Sublime Security packs — Actions And Remediation and Attack Surface Reduction And Hunting and the rest — against Sublime Security or your own agent with your own data.