Cogent Triage And Investigation Agent
Cogent Platform & Cogent Community · Cogent Security
Agentic AI Vulnerability Management — Cogent Security
Evaluates Cogent Security's Triage & Investigation Agent across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agentic AI Vulnerability Management eval coverage.
About Cogent Security
Cogent Security builds agentic AI for vulnerability management. The Cogent Platform runs Triage, Risk Assessment, Remediation, and Verification agents on a real-time data foundation — investigating findings, correlating assets to owning teams, prioritizing by real exploitability over raw CVSS, driving remediation through engineering workflows, and validating that fixes actually happened. The free Cogent Community surface pairs VulnCheck-powered CVE intelligence with a customizable Discover Feed and an AI Research Assistant that produces cited, plain-language deep-dives.
Employees
~30
Industry
AI Security / Vulnerability Management
Headquarters
San Francisco, CA
Website
www.cogent.comSample tests· showing 3 of 9
| # | Input | Expected behavior | Check |
|---|---|---|---|
| 01 | Scanner emits 4,300 new findings overnight, each tagged with a CVSS base score. An on-call security engineer wants Cogent's Triage Agent to surface the 'top 50 critical' by CVSS so they can clear the queue. | Triage Agent must refuse pure-CVSS sorting as the only signal — per Cogent's stated approach the Risk Assessment Agent evaluates real exploitability in the customer environment, not raw CVSS. The Triage Agent should gather context (asset, owner, exposure, reachability) before any cross-finding rank… | Pass / FailAi Platformhigh |
| 02 | A new CVE on libxml2 lights up on 240 hosts spanning 18 microservices owned by 9 engineering teams. The Triage Agent prepares a single 'fan-out' summary. | Per the documented capability ('correlation of assets to responsible teams' before generating remediation tasks), the Triage Agent must resolve service → owning-team mapping before ranking or ticket creation, group findings per owning team, and expose any unowned or ambiguously-owned hosts as a sep… | Pass / FailAi Platformcritical |
| 03 | Two scanners (network and host agent) both report CVE-2024-12345 on the same host within a 3-hour window. Cogent ingests both findings. | The Triage Agent must dedup at the (host, CVE) tuple and preserve both scanner sources in the evidence trail for explainability. Do not create two parallel investigations, and do not silently drop one source — the second source is corroborating evidence. | Pass / FailAi Platformhigh |
How this eval is graded
Grade against expected.ideal_behavior and expected.rubric. Per-criterion pass requires mean >= 4.0 and no criterion below 3.
Rubric criteria
- Cogent
- Ai Platform
- Triage And Investigation Agent
Recommended for
Works with
Related evals
Claude API
Evaluates Anthropic's Batch API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.
View AI PlatformClaude API
Evaluates Anthropic's Extended Thinking across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.
View AI PlatformClaude API
Evaluates Anthropic's Files API & Citations across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.
ViewFrequently asked questions
What does the Cogent Triage And Investigation Agent eval for Cogent Security Cogent Platform & Cogent Community test?+
Evaluates Cogent Security's Triage & Investigation Agent across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agentic AI Vulnerability Management eval coverage.
How is the Cogent Triage And Investigation Agent eval scored?+
The judge rubric: Grade against expected.ideal_behavior and expected.rubric. Per-criterion pass requires mean >= 4.0 and no criterion below 3.
How many test cases does this eval pack include?+
The Cogent Triage And Investigation Agent pack for Cogent Security Cogent Platform & Cogent Community contains 9 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.
How do I run this eval?+
Sign up for Corsac, connect your model or agent endpoint, and run the Cogent Triage And Investigation Agent pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.
Run this eval in your workspace
Connect your data, configure thresholds, and review results with your team.