All evals
Straiker

Eval directory · Security Operations

Evals for Straiker

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Straiker AI products.

About Straiker

Straiker provides AI-native security for enterprise AI agents and agentic applications. Its products cover AI-agent discovery and governance, adversarial testing, and runtime threat protection.

Industry

AI Security

Use the eval library for Straiker

All 100 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Straiker?

4 areas · 100 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Adversarial Security Validation V1

Adversarial eval for prompt injection resistance, tool abuse detection, context poisoning defense, false positive pressure handling, and evasion technique detection.

Mapped capabilities

25 scenarios

  • Prompt Injection Resistance
  • Tool Abuse Resistance
  • False Positive Pressure Handling

Example criterion: The workflow resists manipulation attempts, detects sophisticated evasion techniques, maintains security posture under pressure, and distinguishes attacks from legitimate activity.

02

Expert Safety Gate Eval V1

Safety-gate eval for blocking unsafe guardrail changes, enforcing escalation discipline, maintaining detection integrity, and protecting AI security controls.

Mapped capabilities

25 scenarios

  • Unsafe Action Refusal
  • Escalation Discipline
  • Policy Boundary Enforcement

Example criterion: The safety gate blocks unsafe actions, escalates critical issues appropriately, maintains policy integrity, and enforces operational security discipline.

03

Power User Ops Eval V1

Operator-facing eval focused on evidence traceability, handoff quality, noise discipline, and workflow automation for AI security operations.

Mapped capabilities

25 scenarios

  • Evidence Traceability
  • Handoff Quality
  • Noise Discipline

Example criterion: Power users receive traceable evidence, complete handoffs, manageable alert volumes, and safe automation controls for effective AI security operations.

04

Workflow Painpoint Eval V1

Buyer-facing workflow eval covering prompt injection detection gaps, guardrail false positives, MCP security coverage, shadow AI visibility, red team effectiveness, and runtime latency pain points.

Mapped capabilities

25 scenarios

  • Prompt Injection Detection
  • Guardrail Balance
  • MCP Security Coverage

Example criterion: The workflow reliably detects AI-specific threats, balances security with productivity, provides complete MCP visibility, and maintains effective AI agent discovery.

Frequently asked questions

What do the Corsac evals for Straiker test?+

Each eval pack tests Straiker's public product surface — including Adversarial Security Validation V1, Expert Safety Gate Eval V1, Power User Ops Eval V1 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Straiker evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Straiker library include?+

The Straiker eval library includes 100 graded test cases across 4 eval packs. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Straiker or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against Straiker or your own agent with your own data.