All evals
Harness

Eval directory

Evals for Harness

Eval coverage for Harness, mapped from its public product surface.

About Harness

Harness is a modular AI-native software delivery platform spanning CI/CD, GitOps, internal developer portal, infrastructure as code, and database DevOps. It extends beyond delivery into AI SRE for incident response, AI Security for discovering and protecting AI/MCP assets, and AI Evals for scoring and gating AI agent releases. Pricing is tiered across a Free plan, Essentials, and Enterprise, with an open source offering alongside.

Industry

AI-native DevOps and software delivery platform

Use the eval library for Harness

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Harness?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Continuous Delivery, GitOps & CI

The core delivery path: building artifacts and deploying them to multi-cloud, multi-region, multi-service targets via GitOps or push-based pipelines, including advanced strategies and automated rollback.

Eval suites run as a native pipeline step alongside build, test, and deploy, and gate the release. www.harness.io

Mapped capabilities

4 capabilities

  • Pipeline authoring and build execution

    CI across languages, source providers, and operating systems; pipeline generation from an existing codebase

  • GitOps vs. push-based deployment models

    Choosing and configuring deployment mode for multi-cloud/multi-region/multi-service rollouts

  • Advanced deployment strategies and rollback

    Canary/blue-green style strategies and automated rollback behavior on failure

  • AI-powered deployment verification

    Post-deploy verification signals that inform promotion or rollback decisions

02

Developer Self-Service & Platform Governance

The surfaces that let developers provision and onboard without filing tickets — an enterprise IDP on Backstage, IaC management, and artifact storage — plus the module/plan boundaries that determine what a team can actually use.

Accelerate developer onboarding from months to hours with a low effort enterprise-grade IDP built on top of Backstage. www.harness.io

Mapped capabilities

4 capabilities

  • Internal Developer Portal onboarding

    Backstage-based IDP used to compress developer onboarding from months to hours

  • Infrastructure as Code management

    Standardized provisioning, collaboration, error reduction, and cost control over IaC

  • Artifact registry usage in the delivery path

    Storing and consuming built artifacts, including containerized apps pushed from CI

  • Plan and module entitlement boundaries

    Free, Essentials, Enterprise, and Harness Open Source; modular per-module selection

03

Database DevOps

Bringing schema changes into the same pipeline as application changes, with AI-assisted migration authoring for non-experts and governance controls for DBAs.

Mapped capabilities

4 capabilities

  • Migrations inside the app deployment pipeline

    Deploying database and application changes together rather than out of band

  • AI-assisted migration authoring

    Helping non-SQL-expert developers write complex migrations with database best practices

  • Cross-environment schema visibility

    Comparing schemas between environments to understand change impact and rollout progress

  • Database change governance

    Review/approval of schema changes, policy as code, environment-aware RBAC, audit trails

04

AI Evals & Release Gating

The quality gate for AI agents: scoring responses against metrics and blocking releases that regress, with the same scoring logic applied pre-deploy and against production traffic.

Mapped capabilities

4 capabilities

  • Offline evaluation against golden datasets

    Pre-deploy scoring, prompt and model variant comparison, threshold-based deploy blocking

  • Online evaluation of production traffic

    Scoring real traffic with the same metrics and promoting real failures back into datasets

  • Metric selection and scoring methods

    Built-in and custom metrics via LLM-as-judge, deterministic, and custom methods (e.g. faithfulness, safety, task completion)

  • Eval suite as a native pipeline step

    Running eval suites alongside build/test/deploy so sign-off becomes an automated pass-rate gate

Illustrative example

Input
Our support agent's faithfulness score dropped after a prompt change. How do I keep this version from reaching production?
Expected behavior
Describes running the eval suite as a native pipeline step against a golden dataset before deploy, with a score threshold that fails the step and blocks the release. May note that the same metrics score production traffic online.

05

AI SRE & Incident Response

Incident handling from alert to resolution: correlating signals to recent changes, standardizing first response, routing to the right человек on call, and keeping an authoritative record.

Correlates incident signals with change events from CI/CD, feature flags, infrastructure, and third-party systems. www.harness.io

Mapped capabilities

4 capabilities

  • Automated triage and change correlation

    Correlating incident signals with CI/CD, feature flag, infrastructure, and third-party change events

  • Root cause and blast radius surfacing

    Using change context to surface probable root cause and affected scope

  • Automation runbooks and one-click remediation

    Chained actions — Slack post, Jira ticket, Harness pipeline call, status update, rollback — triggered manually, by rule, or from AI recommendation

  • On-call scheduling and escalation

    Schedules, rotations, escalation policies, and alert routing to the right responder

Illustrative example

Input
Checkout latency alerted overnight. What happens before I'm paged, and where does the incident timeline come from?
Expected behavior
Explains that triage is automated by correlating the alert with recent CI/CD, feature flag, and infrastructure change events to surface probable root cause and blast radius, that on-call schedules route and escalate the page, and that AI Scribe captures the record.

06

AI Security & Agentic Asset Protection

Discovering and defending the AI attack surface — LLMs, MCP servers and tools, AI APIs, and third-party AI services — built on runtime API traffic visibility.

Mapped capabilities

4 capabilities

  • AI and MCP asset discovery

    Continuously updated inventory of every LLM, MCP server/tool/resource, AI API, and third-party AI service, first- and third-party

  • Per-asset risk scoring

    Scoring on authentication status, encryption, internet exposure, sensitive data flows, and vulnerabilities

  • Pre-deployment AI threat testing

    Dynamically testing AI-native applications for prompt injection and data exfiltration before release

  • Runtime attack detection and blocking

    Detecting and blocking live attacks against AI assets in production via AI API monitoring

Coverage is mapped from Harness's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Harness test?+

The coverage map is generated from Harness's own public product surface (AI-native DevOps and software delivery platform): 6 scoring areas — Continuous Delivery, GitOps & CI, Developer Self-Service & Platform Governance, and Database DevOps, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Harness evals scored?+

Every case generated for Harness — across Continuous Delivery, GitOps & CI and Developer Self-Service & Platform Governance and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Harness library include?+

The full Harness library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Pipeline authoring and build execution and GitOps vs. push-based deployment models under Continuous Delivery, GitOps & CI); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Harness or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Harness areas and set them up in a Corsac workspace, where you can run every test case against Harness or your own agent with your own data.