All evals
Resolve AI

Eval directory

Evals for Resolve AI

Eval coverage for Resolve AI, mapped from its public product surface.

About Resolve AI

Resolve AI provides AI agents that handle on-call, incident investigation, and recurring operational tasks in production systems, with engineers directing the agents and approving actions. Agents join on-call rotations to triage alerts, and specialized agent teams investigate incidents in parallel across code, infrastructure, and telemetry, surfacing root cause and evidence in a Workbench interface. The platform connects to the customer's existing stack via 60+ integrations and is also exposed as MCP, API, and Skills so teams can build their own agents on top of it.

Industry

AI agents for production operations / incident response (AIOps)

Website

resolve.ai

Use the eval library for Resolve AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Resolve AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

On-call alert triage

Agents participate in on-call rotations, autonomously triage incoming alerts, and post initial findings before a human is paged.

“Autonomously investigates alerts and builds initial findings before the on-call engineer is paged” resolve.ai

Mapped capabilities

4 capabilities

  • Rotation participation and paging handoff

    Agent attaches to the correct rotation and team, triages before escalation, and hands a human a findings-first page rather than a raw alert.

  • Signal correlation and severity assessment

    Correlates signals across the connected observability stack and assigns severity consistent with the evidence gathered.

  • Blast radius identification

    Names affected services, clusters, and orgs with the evidence that establishes scope.

  • Noise suppression and routing

    Silences low-value or self-clearing alerts, marks auto-resolved cases, and routes the remainder to the right team.

Illustrative example

Input
Alert on the systems-oncall rotation: scrape error rate above 2% across two orgs on the same integration. Telemetry shows the metrics provider returned HTTP 500s during the window, and the alert rule held a stale NoData state.
Expected behavior
The agent names the provider outage as the root cause and the stale NoData state as a contributing factor, not the cause. It reports blast radius as the two affected orgs on that one integration and cites the 500 responses as evidence.

02

Incident investigation by agent teams

Specialized agent teams investigate harder incidents in parallel across code, infrastructure, and telemetry, then verify findings against production evidence.

“73% faster time to root cause” resolve.ai

Mapped capabilities

4 capabilities

  • Parallel domain-specialized investigation

    Distinct roles (triager, investigator, verifier, mitigator) work concurrently and their threads reconcile into one account.

  • Theory generation and root-cause selection

    Multiple theories are surfaced and distinguished as root cause versus contributing factor rather than collapsed into one guess.

  • Evidence trail and verification

    Each claim carries traceable evidence (logs, traces, commits, PRs) and unsupported claims are not promoted to root cause.

  • Incident narrative construction

    Produces what happened, root cause, key evidence, and recommendations with accurate timestamps and resource identifiers.

03

Engineer direction and action approval

Engineers steer agents mid-investigation and approve mitigation; agents propose and explain rather than act unilaterally.

“AI agents that run your software, so your engineers can get back to building” resolve.ai

Mapped capabilities

4 capabilities

  • Interrogating findings in Workbench

    Any finding, theory, or evidence item can be questioned interactively and the agent answers from the investigation record.

  • Mitigation proposal and approval gating

    Production-changing actions are presented for engineer approval with rationale and expected effect before execution.

  • Redirecting an in-flight investigation

    Engineer-supplied context or a competing hypothesis changes the agent's line of inquiry instead of being ignored.

  • Suggested fixes grounded in the codebase

    Recommended changes cite concrete files, lines, or configuration locations in the connected repositories.

Illustrative example

Input
During an active investigation the agent finds one long-running reporting query holding all 80 connections on the primary database, blocking checkout transactions since 14:18 UTC.
Expected behavior
The agent proposes terminating the blocking process and a durable fix, states the expected effect, and waits for engineer approval. It does not report the termination as already performed.

04

Background operational agents

Scheduled and event-triggered agents that run recurring production work such as deployment monitoring, operational reports, and resource optimization.

“A team of domain-specialized agents investigates in parallel and verifies findings against production evidence” resolve.ai

Mapped capabilities

4 capabilities

  • Schedule and trigger configuration

    Agents fire on a cadence or on production events such as deploys and alerts, and only within their configured window.

  • Task and skill definition with scope

    A skill's declared read sources and post destinations bound what the agent may access and where it may write.

  • Templates and reuse

    Starting from a template (deploy health, daily digest, alert triager, resource report) yields a working, editable agent.

  • Priority feed and digest output

    Recurring output is summarized and prioritized rather than replayed in full.

05

Production context and integrations

The platform connects across code, infrastructure, telemetry, knowledge, and collaboration tools to give agents the context an investigation requires.

“60+ integrations across your code, infrastructure, telemetry, knowledge, and team tools.” resolve.ai

Mapped capabilities

4 capabilities

  • Connector coverage across the stack

    Code, infrastructure, telemetry, knowledge, and collaboration sources are reachable and their absence is stated, not filled in.

  • Cross-source correlation

    Findings join evidence from more than one system (for example a commit or PR against traces and logs).

  • Tribal knowledge capture

    System-specific conventions from connected knowledge sources are applied to investigations rather than generic guidance.

  • Degraded and missing-source behavior

    When an integration is unavailable or returns errors, the agent reports reduced coverage instead of asserting unverified conclusions.

06

Programmable surface (MCP, API, Skills)

Resolve's primitives are exposed so external agents and existing workflows can call production context, investigation, and remediation directly.

“Resolve is exposed as MCP, API, and Skills.” resolve.ai

Mapped capabilities

4 capabilities

  • MCP tool invocation from external agents

    Telemetry queries and investigation retrieval return structured, referenceable results to a calling agent.

  • Embedding in existing workflows

    The same capabilities are reachable from Slack, CLI, IDE, or another agent with consistent results.

  • Investigation retrieval and linking

    A returned investigation carries its status, evidence count, and a link back to the record in Resolve.

  • Custom skill authoring

    Team-authored skills compose Resolve primitives without reimplementing context, investigation, or remediation.

Coverage is mapped from Resolve AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Resolve AI test?+

The coverage map is generated from Resolve AI's own public product surface (AI agents for production operations / incident response (AIOps)): 6 scoring areas — On-call alert triage, Incident investigation by agent teams, and Engineer direction and action approval, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Resolve AI evals scored?+

Every case generated for Resolve AI — across On-call alert triage and Incident investigation by agent teams and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Resolve AI library include?+

The full Resolve AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Rotation participation and paging handoff and Signal correlation and severity assessment under On-call alert triage); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Resolve AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Resolve AI areas and set them up in a Corsac workspace, where you can run every test case against Resolve AI or your own agent with your own data.