All evals
CS

Eval directory · AI Platform

Evals for Cogent Security

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Cogent Security AI products.

About Cogent Security

Cogent Security builds agentic AI for vulnerability management. The Cogent Platform runs Triage, Risk Assessment, Remediation, and Verification agents on a real-time data foundation — investigating findings, correlating assets to owning teams, prioritizing by real exploitability over raw CVSS, driving remediation through engineering workflows, and validating that fixes actually happened. The free Cogent Community surface pairs VulnCheck-powered CVE intelligence with a customizable Discover Feed and an AI Research Assistant that produces cited, plain-language deep-dives.

Employees

~30

Industry

AI Security / Vulnerability Management

Headquarters

San Francisco, CA

Use the eval library for Cogent Security

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Cogent Security?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Cogent Cogent Community Discover Feed And Research Assistant

Evaluates Cogent Security's Cogent Community: Discover Feed & Research Assistant across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agentic AI Vulnerability Management eval coverage.

Mapped capabilities

9 scenarios

  • Discover Feed customization scope
  • Research Assistant context injection
  • prompt-injection in CVE description

Public sample case

Input
User customizes the Discover Feed to follow only Java ecosystem CVEs. A non-Java CVE that is on KEV with mass exploitation appears.
Expected behavior
Per the documented Discover Feed ('customizable, real-time feed for vulnerability and exploit intelligence showing breaking disclosures, trending activity, and the topics a user follows'), the feed must honor the customization filter for the main stream AND surface a distinct 'breaking disclosure /…
Check
Pass / fail check

02

Cogent Remediation Orchestration And Ticketing

Evaluates Cogent Security's Remediation Orchestration & Ticketing across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agentic AI Vulnerability Management eval coverage.

Mapped capabilities

9 scenarios

  • ticket created per owning team
  • remediation steps explainability
  • ticket idempotency on re-ingest

Public sample case

Input
Cluster of CVE-2024-12345 findings spans 18 services across 9 owning teams. Remediation Agent prepares ticket creation.
Expected behavior
Per the documented capability ('creation of remediation tasks for engineers') and asset-ownership correlation, Remediation Agent must create one ticket per owning team scoped to that team's assets, not a single global ticket and not 240 tickets per host. Each ticket must carry the per-team asset li…
Check
Pass / fail check

03

Cogent Risk Assessment And Contextual Prioritization

Evaluates Cogent Security's Risk Assessment & Contextual Prioritization across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agentic AI Vulnerability Management eval coverage.

Mapped capabilities

9 scenarios

  • exploit-in-the-wild outranks CVSS-critical
  • confidence level surfaced with verdict
  • environmental exposure overrides

Public sample case

Input
Finding A: CVSS 9.8, no public exploit, internal-only host behind two VPN hops. Finding B: CVSS 7.5, KEV-listed with mass-exploitation underway, on an internet-facing edge service.
Expected behavior
Per Cogent's stated principle — agents 'evaluate real exploitability in the environment, not just CVSS severity' — Finding B must rank above Finding A. The explanation must cite KEV listing, mass-exploitation telemetry, and exposure (internet-facing vs internal-only) as the dominant inputs, with CV…
Check
Pass / fail check

04

Cogent Safety Governance And Human In The Loop

Evaluates Cogent Security's Safety, Governance & Human-in-the-Loop across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agentic AI Vulnerability Management eval coverage.

Mapped capabilities

10 scenarios

  • no autonomous remediation without approval
  • agent reasoning audit log
  • model output safety on jailbreak

05

Cogent Scanner Ingestion Asset Graph And Normalization

Evaluates Cogent Security's Scanner Ingestion, Asset Graph & Normalization across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agentic AI Vulnerability Management eval coverage.

Mapped capabilities

9 scenarios

  • cross-scanner CVE id normalization
  • asset identity reconciliation
  • ephemeral asset (container, lambda)

06

Cogent Triage And Investigation Agent

Evaluates Cogent Security's Triage & Investigation Agent across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agentic AI Vulnerability Management eval coverage.

Mapped capabilities

9 scenarios

  • raw CVSS-only triage rejected
  • asset–owner correlation before ticket
  • duplicate finding dedup across scanners

07

Cogent Verification And Closure Validation

Evaluates Cogent Security's Verification & Closure Validation across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agentic AI Vulnerability Management eval coverage.

Mapped capabilities

9 scenarios

  • ticket closed != fix verified
  • patch present but service not restarted
  • config-change remediation verification

08

Cogent Vulncheck Powered Cve Knowledge Base

Evaluates Cogent Security's VulnCheck-Powered CVE Knowledge Base across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agentic AI Vulnerability Management eval coverage.

Mapped capabilities

9 scenarios

  • multi-source citation discipline
  • source disagreement surfaced
  • stale NVD data refresh

Frequently asked questions

What do the Corsac evals for Cogent Security test?+

Each eval pack tests Cogent Security's public product surface — including Cogent Cogent Community Discover Feed And Research Assistant, Cogent Remediation Orchestration And Ticketing, and Cogent Risk Assessment And Contextual Prioritization — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Cogent Security evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Cogent Security cases — from Cogent Safety Governance And Human In The Loop (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Cogent Security library.

How many test cases does the Cogent Security library include?+

The Cogent Security eval library includes 73 graded test cases across 8 eval packs, the largest being Cogent Safety Governance And Human In The Loop with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Cogent Security or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Cogent Security packs — Cogent Cogent Community Discover Feed And Research Assistant and Cogent Remediation Orchestration And Ticketing and the rest — against Cogent Security or your own agent with your own data.