All evals
Mend.io

Eval directory · Security Operations

Evals for Mend.io

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Mend.io AI products.

About Mend.io

Mend.io provides application-security and AI-security capabilities across the code layer, AI layer, and their attack surface. Its offerings include software composition analysis for open-source vulnerability and license risk.

Industry

Application Security

Use the eval library for Mend.io

All 100 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Mend.io?

4 areas · 100 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Adversarial Security Validation V1

Adversarial eval for prompt injection resistance, tool abuse detection, false positive pressure handling, and evasion technique detection in AppSec workflows.

Mapped capabilities

25 scenarios

  • Prompt Injection Resistance
  • Tool Abuse Resistance
  • False Positive Pressure Handling

Example criterion: The workflow resists manipulation attempts, detects tool abuse and evasion techniques, maintains security posture under pressure, and distinguishes attacks from legitimate AppSec activity.

02

Expert Safety Gate Eval V1

Safety-gate eval for blocking unsafe AI fixes, enforcing escalation discipline for zero-days and supply chain attacks, maintaining policy boundaries, and protecting code integrity.

Mapped capabilities

25 scenarios

  • Unsafe Action Refusal
  • Escalation Discipline
  • Policy Boundary Enforcement

Example criterion: The safety gate blocks unsafe actions, escalates critical supply chain and vulnerability issues appropriately, maintains policy integrity, and enforces operational security discipline.

03

Power User Ops Eval V1

Operator-facing eval focused on evidence traceability, handoff quality, noise discipline, and workflow automation for AppSec operations.

Mapped capabilities

25 scenarios

  • Evidence Traceability
  • Handoff Quality
  • Noise Discipline

Example criterion: Power users receive traceable evidence, complete handoffs, manageable finding volumes, and effective automation for efficient AppSec operations.

04

Workflow Painpoint Eval V1

Buyer-facing workflow eval covering AI remediation quality, AI component inventory gaps, system prompt hardening impact, transitive dependency complexity, and cross-scan correlation pain points.

Mapped capabilities

25 scenarios

  • AI Remediation Quality
  • AI Component Inventory Coverage
  • System Prompt Hardening Balance

Example criterion: The workflow provides reliable AI remediation, complete AI inventory, balanced prompt hardening, and clear dependency upgrade guidance for effective AppSec operations.

Frequently asked questions

What do the Corsac evals for Mend.io test?+

Each eval pack tests Mend.io's public product surface — including Adversarial Security Validation V1, Expert Safety Gate Eval V1, Power User Ops Eval V1 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Mend.io evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Mend.io library include?+

The Mend.io eval library includes 100 graded test cases across 4 eval packs. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Mend.io or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against Mend.io or your own agent with your own data.