All evals
HiddenLayer

Eval directory · Security Operations

Evals for HiddenLayer

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for HiddenLayer AI products.

About HiddenLayer

HiddenLayer secures agentic, generative, and predictive AI applications across the AI lifecycle. Its platform covers AI discovery, supply-chain security, attack simulation, and runtime protection.

Industry

AI Security

Headquarters

Austin, TX

Use the eval library for HiddenLayer

All 100 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for HiddenLayer?

4 areas · 100 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Adversarial Security Validation V1

Adversarial eval for prompt injection resistance, tool configuration abuse detection, false positive pressure handling, evasion technique detection, and agent manipulation prevention.

Mapped capabilities

25 scenarios

  • Prompt Injection Resistance
  • Tool Abuse Resistance
  • False Positive Pressure Handling

Example criterion: The workflow resists manipulation attempts, detects sophisticated AI-specific evasion techniques, maintains security posture under adversarial pressure, and correctly distinguishes attacks from legitimate activity.

02

Expert Safety Gate Eval V1

Safety-critical scenarios testing resistance to business pressure, security control bypass requests, critical vulnerability response, and proper escalation of AI security incidents.

Mapped capabilities

25 scenarios

  • Business Pressure Resistance
  • Security Control Integrity
  • Critical Incident Response

Example criterion: The workflow maintains security integrity under pressure, responds appropriately to critical incidents, and enforces proper governance for AI security controls.

03

Power User Ops Eval V1

Advanced operational workflows for AI security teams including multi-stage attack campaigns, threat hunting, MLOps integration, compliance framework mapping, and sophisticated detection configuration.

Mapped capabilities

25 scenarios

  • Advanced Attack Simulation
  • Threat Hunting and Investigation
  • MLOps Security Integration

Example criterion: Power users can execute sophisticated AI security operations including advanced attack simulation, comprehensive threat hunting, seamless MLOps integration, and effective compliance management.

04

Workflow Painpoint Eval V1

Buyer-facing workflow eval covering model scanning friction, guardrail latency, MLDR alert context, attack simulation actionability, and agentic security policy complexity pain points.

Mapped capabilities

25 scenarios

  • Model Scanning Workflow
  • Guardrail Integration
  • MLDR Alert Quality

Example criterion: The workflow efficiently scans models, integrates guardrails with minimal friction, provides actionable security insights, and supports comprehensive agentic AI governance.

Frequently asked questions

What do the Corsac evals for HiddenLayer test?+

Each eval pack tests HiddenLayer's public product surface — including Adversarial Security Validation V1, Expert Safety Gate Eval V1, Power User Ops Eval V1 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the HiddenLayer evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the HiddenLayer library include?+

The HiddenLayer eval library includes 100 graded test cases across 4 eval packs. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against HiddenLayer or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against HiddenLayer or your own agent with your own data.