All evals
Cleric

Eval directory

Evals for Cleric

Eval coverage for Cleric, mapped from its public product surface.

About Cleric

Cleric is an AI site reliability engineering agent that follows changes into production, investigates incidents, proposes fixes, and learns from outcomes. It picks up work from change regressions, alerts, support tickets, scheduled checks, and engineer requests, and verifies deployed changes against real traffic. It integrates with tools like Datadog, PagerDuty, GitHub, and Slack, and is billed per resolved Issue through a monthly credit pool rather than per token.

Industry

AI SRE / production incident investigation agent

Website

cleric.ai

Use the eval library for Cleric

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Cleric?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Change Verification

Following a pull request from open through deploy and checking that production behaves as expected against real traffic, rather than waiting for an alert.

“Cleric tracks each PR from open to deployed, then checks that production behaves as expected.” cleric.ai

Mapped capabilities

4 capabilities

  • PR-to-production tracking

    Associating an open PR with the deployed revision actually serving traffic.

  • Expected-outcome recording

    Capturing what the change is supposed to do at PR time and checking against it later.

  • Verification window handling

    Reporting in-progress verification state, elapsed checks, and next run rather than premature all-clear.

  • Regression detection on real traffic

    Distinguishing no-regression-detected from insufficient-signal after a change reaches production.

Illustrative example

Input
PR #1849 (checkout-api: tighten timeout handling) merged 40 minutes ago and reached production. An engineer asks in Slack whether the change is safe to leave running.
Expected behavior
Reports the change as still inside its verification window, notes that production checks are currently passing against real traffic, and gives the next check time — without declaring the change verified or the Issue resolved.

02

Incident Investigation and Root Cause

Opening and running an investigation that tests possible causes against production systems and returns findings an engineer can act on.

“5 min Time to Root Cause 92% Actionable Findings 200,000+ Production-Grade Investigations” cleric.ai

Mapped capabilities

4 capabilities

  • Multi-trigger intake

    Same investigation flow from change regression, alert, support ticket, scheduled check, or engineer request.

  • Alert grouping by shared cause

    Collapsing many correlated alerts into a single investigation instead of parallel duplicates.

  • Hypothesis testing against production

    Probing candidate causes with live systems rather than asserting a cause from alert text.

  • Actionable finding quality

    Findings that name affected services, evidence, and next step for the on-call engineer.

Illustrative example

Input
Twenty-three alerts fire within four minutes across payment-api and two downstream services shortly after a deploy. The on-call engineer asks Cleric what is happening.
Expected behavior
Opens a single investigation that groups the correlated alerts under one suspected cause, names the affected services and the triggering change, and reports the group count instead of returning separate per-alert investigations.

03

Fix Proposal and Outcome Verification

Proposing a remediation and then confirming against production whether the problem is actually fixed, treating diagnosis and resolution as separate bars.

Mapped capabilities

4 capabilities

  • Proposed fix grounded in findings

    Remediation that follows from the investigated cause, not a generic mitigation.

  • Diagnosis vs. resolution separation

    Not claiming a fix landed when only a cause was identified.

  • Post-fix verification against production

    Checking the deployed fix against live behavior before closing.

  • Issue closure criteria

    Closing only the work an engineer would have considered closed.

04

Operational Memory and Decision Transparency

Persisting engineering judgment across investigations and exposing the agent's decisions, not just its outputs, for human review.

Mapped capabilities

4 capabilities

  • Reuse of prior investigation knowledge

    Applying previously verified operational context to a recurring problem.

  • Learning from outcomes

    Incorporating whether a past fix held when handling a similar signal.

  • Reviewable decision trail

    Surfacing why a path was taken so a reviewer can audit judgment.

  • Investigation quality regression checks

    Comparative ranking of investigation traces to catch quality regressions.

05

Integrations and Engineer Workflow Surfaces

Where Cleric picks up work and hands it back: observability, paging, source control, chat, and configurable agent entry points.

“You pay when Cleric closes the Issue an engineer would have closed.” cleric.ai

Mapped capabilities

4 capabilities

  • Observability and paging tools

    Datadog and PagerDuty as investigation signal and context sources.

  • GitHub and Slack handoff

    Change context from GitHub; findings and requests exchanged in Slack.

  • Scheduled Monitors and Custom Agents

    Recurring checks and team-defined agents as first-class triggers.

  • Slash Commands

    Out-of-band engineer requests opening the standard investigation flow.

06

Access Controls and Commercial Model

How Cleric reaches production safely and how work is metered, covering the questions security reviewers and budget owners raise before a pilot.

“Agents that investigate and fix production problems before and after they page you.” cleric.ai

Mapped capabilities

4 capabilities

  • Read-only default posture

    Default non-mutating access to production systems.

  • Private and VPC resource access

    Reaching internal resources over controlled network paths.

  • Enterprise identity and audit

    SSO, OIDC, SAML, and audit logs for organization-wide rollout.

  • Credit accounting per resolved Issue

    Charging per closed Issue or Custom Work against a monthly credit pool, not per token.

Coverage is mapped from Cleric's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Cleric test?+

The coverage map is generated from Cleric's own public product surface (AI SRE / production incident investigation agent): 6 scoring areas — Change Verification, Incident Investigation and Root Cause, and Fix Proposal and Outcome Verification, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Cleric evals scored?+

Every case generated for Cleric — across Change Verification and Incident Investigation and Root Cause and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Cleric library include?+

The full Cleric library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, PR-to-production tracking and Expected-outcome recording under Change Verification); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Cleric or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Cleric areas and set them up in a Corsac workspace, where you can run every test case against Cleric or your own agent with your own data.