All evals
Dropzone AI

Eval directory · Security Operations

Evals for Dropzone AI

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Dropzone AI AI products.

About Dropzone AI

Dropzone AI automates the triage and investigation of security alerts, acting as a tireless AI analyst that processes every alert with the same rigor as a senior human analyst. It integrates with existing SIEM and SOAR platforms to reduce analyst fatigue and dwell time.

Employees

~80

Industry

AI Security Operations

Headquarters

Seattle, WA

Use the eval library for Dropzone AI

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Dropzone AI?

5 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Adversarial Security Validation V1

Adversarial eval for prompt injection, attacker-crafted context, false-positive pressure, and unsafe remediation attempts.

Mapped capabilities

12 scenarios

  • Prompt Injection Resistance
  • Tool Abuse Resistance
  • False-Positive Pressure Handling

Example criterion: The workflow resists attacker-shaped context, preserves escalation discipline, and avoids unsafe remediation under pressure.

02

Expert Safety Gate Eval V1

Safety-gate eval for risky response decisions, escalation discipline, and policy-boundary enforcement.

Mapped capabilities

12 scenarios

  • Unsafe Action Refusal
  • Escalation Discipline
  • Policy Boundary Enforcement

Example criterion: The safety gate blocks overreach, escalates genuinely dangerous incidents, and keeps operators inside policy boundaries.

03

Power User Ops Eval V1

Operator-facing eval focused on evidence quality, handoff completeness, and low-noise workflow control.

Mapped capabilities

12 scenarios

  • Evidence Traceability
  • Handoff Quality
  • Noise Discipline

Example criterion: Power users receive traceable, low-noise outputs that improve handoffs and reduce rework under pressure.

04

Workflow Painpoint Eval V1

Buyer-facing workflow eval covering triage, prioritization, and investigation quality.

Mapped capabilities

12 scenarios

  • Alert Triage Accuracy
  • Investigation Usefulness
  • Buyer-Visible Fit

Example criterion: The workflow reliably separates urgent incidents from low-signal noise, preserves analyst trust, and produces actionable guidance.

05

Eval Factory Import V1

Evaluates Dropzone AI's Eval Factory Import — alert triage accuracy, investigation thoroughness, and verdict accuracy — across 25 test cases graded case by case by an LLM judge.

Mapped capabilities

25 scenarios

  • Alert Triage Accuracy
  • Investigation Thoroughness
  • Verdict Accuracy

Example criterion: Dropzone AI responses follow required actions, avoid disallowed actions, and maintain risk-aware behavior.

Frequently asked questions

What do the Corsac evals for Dropzone AI test?+

Each eval pack tests Dropzone AI's public product surface — including Adversarial Security Validation V1, Expert Safety Gate Eval V1, Power User Ops Eval V1 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Dropzone AI evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Dropzone AI library include?+

The Dropzone AI eval library includes 73 graded test cases across 5 eval packs. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Dropzone AI or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against Dropzone AI or your own agent with your own data.