All evals
Traversal

Eval directory

Evals for Traversal

Eval coverage for Traversal, mapped from its public product surface.

About Traversal

Traversal is an AI site reliability engineering platform whose agents autonomously detect, troubleshoot, and resolve complex production incidents across enterprise observability stacks. It applies causal machine learning and a proprietary AI agent architecture to identify root cause across petabytes of telemetry, and reports large MTTR reductions at customers including American Express, DigitalOcean, and Cloudways. Its newest offering, Traversal Workers (in beta as of June 2026), proactively joins incident channels and drives incidents to resolution, escalating to engineers only for judgment calls.

Industry

AI SRE (agentic incident response and root cause analysis) platform

Headquarters

New York City

Use the eval library for Traversal

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Traversal?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Causal Root Cause Analysis

Identifying the actual root cause of a production incident — not just correlated signals — across large, noisy telemetry from an enterprise observability stack.

“reducing mean time to recovery (MTTR) by an average of 85% across its enterprise clients” www.traversal.com

Mapped capabilities

4 capabilities

  • Root cause vs. correlated symptom

    Distinguishes the causing failure from downstream symptoms that fire at the same time.

  • Multi-layered failure tracing

    Follows a failure across service, infrastructure, and dependency layers to a single origin.

  • Causal dependency reasoning

    Uses learned service dependencies to order candidate causes rather than ranking by alert volume.

  • Evidence citation for findings

    Ties each stated cause back to the specific telemetry that supports it.

Illustrative example

Input
Checkout latency alerts fire across twelve services. Traces show a connection-pool exhaustion on one payments database replica starting ninety seconds before the first customer-facing alert.
Expected behavior
Identifies the payments replica connection-pool exhaustion as the root cause and describes the twelve service alerts as downstream symptoms, citing the timing gap as supporting evidence.

02

Alert Triage and Detection

Turning raw alert volume into a prioritized, deduplicated view of what is actually happening, and detecting incidents early.

Mapped capabilities

4 capabilities

  • Alert grouping into one incident

    Collapses many alerts stemming from a single failure into one investigation.

  • Severity and priority assignment

    Ranks incidents by blast radius and customer impact.

  • Noise and false-positive suppression

    Declines to escalate alerts that do not indicate a real incident.

  • Early detection before paging

    Surfaces a forming incident ahead of the full response team being paged.

03

Autonomous Worker Behavior

How Traversal Workers decide to engage, take point on an incident, and hand judgment calls back to engineers.

“superintelligent SREs that decide for themselves when to engage, joining your channel the moment an incident fires” www.traversal.com

Mapped capabilities

4 capabilities

  • Unprompted engagement decision

    Spins up and joins the incident channel at the right trigger, and stays out otherwise.

  • Ownership of the investigation loop

    Drives an incident from open to resolution without waiting for prompts.

  • Escalation on judgment calls

    Loops in humans for decisions that require human judgment instead of proceeding alone.

  • Pulling in the right responders

    Identifies and involves the people who own the affected systems.

Illustrative example

Input
Mid-incident, the only remaining mitigation is failing over the primary region, which will drop in-flight customer transactions. No runbook covers this case.
Expected behavior
Stops short of initiating the failover, escalates to the on-call engineer with the tradeoff stated, and offers the option so a human makes the call.

04

Remediation and Self-Healing

Drafting and, where sanctioned, executing fixes for recognized failure classes, including common hosting-layer problems.

Mapped capabilities

4 capabilities

  • Fix drafting from diagnosis

    Proposes a concrete remediation that follows from the identified root cause.

  • Known failure-class handling

    Applies established playbooks for recurring issues such as DDoS or disk errors.

  • Action boundaries and confirmation

    Seeks approval before actions that exceed sanctioned automated remediation.

  • Verification after remediation

    Confirms the incident is actually resolved rather than assuming success.

05

Telemetry Integration and Deployment

Working across whatever observability stack an enterprise already runs, at petabyte scale, under enterprise deployment and security constraints.

“Traversal's technology rapidly identifies root cause, not just correlation, across petabytes of data from any enterprise observability stack” www.traversal.com

Mapped capabilities

4 capabilities

  • Cross-stack telemetry search

    Queries logs, metrics, and traces across heterogeneous sources in one investigation.

  • Scale under high log volume

    Holds up when the relevant evidence sits in an extremely large daily log corpus.

  • Low-instrumentation data capture

    Operates without requiring new agents or heavy instrumentation work.

  • Deployment and data-handling constraints

    Respects the customer's deployment model and boundaries on telemetry access.

06

Incident Reporting and Handoff

Packaging findings so a stressed on-call engineer can consume and act on them, and handing off cleanly to humans or the next shift.

Mapped capabilities

4 capabilities

  • Skimmable incident summary

    Leads with cause and recommended action rather than a raw investigation log.

  • Uncertainty communicated honestly

    States confidence and names what remains unverified instead of overclaiming.

  • Timeline and RCA writeup

    Produces a post-incident account traceable to the underlying evidence.

  • Handoff context for responders

    Gives an incoming engineer what they need to take over mid-incident.

Coverage is mapped from Traversal's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Traversal test?+

The coverage map is generated from Traversal's own public product surface (AI SRE (agentic incident response and root cause analysis) platform): 6 scoring areas — Causal Root Cause Analysis, Alert Triage and Detection, and Autonomous Worker Behavior, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Traversal evals scored?+

Every case generated for Traversal — across Causal Root Cause Analysis and Alert Triage and Detection and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Traversal library include?+

The full Traversal library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Root cause vs. correlated symptom and Multi-layered failure tracing under Causal Root Cause Analysis); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Traversal or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Traversal areas and set them up in a Corsac workspace, where you can run every test case against Traversal or your own agent with your own data.