All evals
7A

Eval directory

Evals for 7AI

Eval coverage for 7AI, mapped from its public product surface.

About 7AI

7AI is an agentic security platform that uses AI agents to take on the repetitive, non-human work of the security operations center, with humans on the loop. Its agents detect, investigate, respond, and hunt — covering threat hunting, alert enrichment, false positive elimination, and end-to-end autonomous investigations. Founded in 2024 by Cybereason co-founders Lior Div and Yonatan Striem-Amit, it came out of stealth in February 2025.

Industry

agentic AI security operations (SOC) platform

Use the eval library for 7AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for 7AI?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Alert Triage and False Positive Elimination

The core queue work 7AI advertises: taking alerts from identity, cloud, and EDR sources, adding context, and separating noise from what deserves a human.

Eliminating noise while ensuring alerts get attention. 7ai.com

Mapped capabilities

4 capabilities

  • Alert enrichment from identity, cloud, and EDR telemetry

    Gathering and attaching the surrounding context an analyst would otherwise assemble by hand.

  • False positive disposition with stated rationale

    Closing benign alerts while making the reasoning inspectable rather than opaque.

  • Noise reduction that preserves true positives

    Behavior on alerts that superficially resemble known-benign patterns but are not.

  • Escalation routing to a human analyst

    When and how an alert is handed up rather than auto-resolved.

Illustrative example

Input
An EDR alert fires for encoded PowerShell on a finance workstation, launched by a parent process listed on the customer's documented software-deployment allowlist.
Expected behavior
The agent enriches the alert with endpoint and identity context, dispositions it as a false positive, and names the allowlist match as the reason. It closes the alert without escalating to a human analyst.

02

Autonomous End-to-End Investigation

7AI's demo promises autonomous investigations that show every step along the way; this area covers the investigation itself and the legibility of its trace.

Autonomous AI agents run full investigations, showing every step along the way. 7ai.com

Mapped capabilities

4 capabilities

  • Multi-step investigation execution to a conclusion

    Driving an alert through pivots to a stated verdict without human prompting at each step.

  • Step-by-step reasoning trace

    Exposing the sequence of actions and queries taken during the investigation.

  • Evidence-linked findings

    Tying each conclusion back to specific observed telemetry rather than assertion.

  • Scope and blast-radius determination

    Identifying which other hosts, identities, or assets are implicated.

03

Threat Hunting and Intel Application

Covers Threat Hunt and Threat Intel Hunt: ingesting threat data, understanding IOCs, and determining whether they exist in the customer environment.

How 7AI agents ingest threat data, understand IOCs, and determine whether they exist in your environment. 7ai.com

Mapped capabilities

4 capabilities

  • IOC ingestion and normalization by type

    Parsing hashes, domains, IPs, and other indicators out of supplied threat data.

  • Environment presence determination

    Answering whether an indicator is actually observed in the customer's telemetry.

  • Intel-report-driven hunt construction

    Turning a narrative threat report into an executable hunt.

  • Hunt result reporting and negative findings

    Communicating absence, coverage gaps, and unsearched sources honestly.

Illustrative example

Input
Given a threat intel report containing three indicators — one domain, one SHA-256 hash, one IPv4 address — determine whether any are present in the connected environment.
Expected behavior
The agent classifies each indicator by type, queries the connected telemetry, and returns a separate present-or-absent verdict per indicator, naming the data source and time window searched for each one.

04

Human-on-the-Loop Control

7AI positions humans as on the loop rather than out of it, and ships Skills so teams can make agents work their way. This area covers steerability and handoff.

AI agents that detect, investigate, respond, and hunt, and humans on the loop 7ai.com

Mapped capabilities

4 capabilities

  • Directing agents via Skills

    Applying team-authored instructions that shape how agents work a case.

  • Approval gating on response actions

    Pausing for human confirmation before an agent acts on the environment.

  • Analyst override and correction handling

    Behavior when a human disagrees with an agent disposition.

  • Handoff summaries for escalated cases

    What a receiving human analyst is given when work moves up the funnel.

05

Data Connectivity and Deployment Surface

The substrate the agents run on: connected security data sources, Federated SIEM, and 7AI Build for customer-authored extensions, including partner-led deployments.

Mapped capabilities

4 capabilities

  • Identity, cloud, and EDR source connections

    Reading from the alert sources named in 7AI's demo agenda.

  • Federated SIEM querying

    Reaching across federated data without centralizing it first.

  • Custom agent and skill authoring via 7AI Build

    Letting teams extend agent behavior for their own environment.

  • Multi-customer and partner deployment

    Operating across tenants in the channel and alliance model 7AI launched.

06

Outcome Reporting and Claim Integrity

7AI leads with quantified outcomes — alerts processed, analyst hours saved, productivity reclaimed. This area covers whether reported outcomes are traceable and honestly bounded.

9M+ Alerts Processed By the 7AI Agentic Security Platform 7ai.com

Mapped capabilities

3 capabilities

  • Analyst-time-saved accounting

    How saved-hours figures are derived and attributed.

  • Alert volume and disposition reporting

    Reconciling processed, auto-closed, and human-reviewed counts.

  • Traceability of customer-facing metrics

    Whether a headline number can be drilled into the underlying cases.

Coverage is mapped from 7AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for 7AI test?+

The coverage map is generated from 7AI's own public product surface (agentic AI security operations (SOC) platform): 6 scoring areas — Alert Triage and False Positive Elimination, Autonomous End-to-End Investigation, and Threat Hunting and Intel Application, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the 7AI evals scored?+

Every case generated for 7AI — across Alert Triage and False Positive Elimination and Autonomous End-to-End Investigation and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the 7AI library include?+

The full 7AI library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Alert enrichment from identity, cloud, and EDR telemetry and False positive disposition with stated rationale under Alert Triage and False Positive Elimination); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against 7AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped 7AI areas and set them up in a Corsac workspace, where you can run every test case against 7AI or your own agent with your own data.