All evals
O

Eval directory

Evals for OpenHands

Eval coverage for OpenHands, mapped from its public product surface.

About OpenHands

OpenHands is an open-source platform for running autonomous AI coding agents locally, in OpenHands Cloud, or self-hosted in a customer VPC. Its Agent Canvas desktop workspace runs multiple agents in parallel, each in its own git worktree, and connects to third-party harnesses like Claude Code, Codex, and Gemini CLI via the Agent Client Protocol. It also offers Slack, GitHub, Jira, and Linear automations plus a terminal UI, web GUI, and SDK for building custom agents.

Industry

open-source cloud coding agent platform

Use the eval library for OpenHands

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for OpenHands?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agent Canvas workspace

The local visual workspace that is the new face of OpenHands: running several agents at once from one interface, with per-agent isolation and configurable extensions.

Run multiple agents simultaneously, each isolated in its own git worktree. www.openhands.dev

Mapped capabilities

4 capabilities

  • Parallel agent sessions

    Running multiple agents simultaneously and viewing all conversations from a single Agent Canvas home.

  • Git worktree isolation

    Each agent isolated in its own git worktree rather than a single shared session.

  • Backend switching

    Switching between local machine, remote VM, and OpenHands Cloud from the header without editing config.

  • Extensions and MCP

    Shipped library of MCP connections and Agent Skills, plus custom additions.

02

Bring-your-own-agent via ACP

Connecting existing agent harnesses through the open Agent Client Protocol so teams keep their current subscriptions, models, and workflows.

Mapped capabilities

4 capabilities

  • Supported harnesses

    Claude Code, Codex, and Gemini CLI listed as supported ACP connections.

  • Subscription and model reuse

    Using an existing vendor subscription and familiar model instead of switching tools.

  • Model-agnostic OpenHands Agent

    The open-source OpenHands Agent with a bring-your-own LLM key.

  • ACP compatibility surface

    Connecting any ACP-compatible harness beyond the named three.

03

Automations and work-tool integrations

Template-driven automations that run where the team already works, with cron, polling, and event-driven triggers across 70+ integrations.

Pick a template, connect Slack, GitHub, Linear, or 70+ integrations www.openhands.dev

Mapped capabilities

4 capabilities

  • Trigger semantics

    Cron, polling, and event-driven triggers, including once-per-label-event behavior for PR review.

  • Chat and ticketing automations

    Slack @openhands channel monitoring, Linear issue classification and duplicate detection, Jira incident triage and routing.

  • Code and CI automations

    GitHub PR review comments, CI failure diagnosis with a proposed-fix PR, security alert remediation PRs.

  • Cross-tool drafting

    Incident retrospective assembly from Notion and Jira into a timeline with owners, decisions, and follow-ups.

Illustrative example

Input
Our GitHub repo uses the OpenHands code review automation on a 'needs-ai-review' label. If a teammate removes the label and adds it again, will the agent post a second review comment?
Expected behavior
Explains that the automation triggers on the configured pull request label and posts one AI review comment per label event, so re-adding the label fires a new review. Does not claim it reviews every push or commit.

04

Deployment and runtime control

Where and how agents actually execute: laptop, hosted cloud, or self-hosted in a customer VPC, with the sandbox runtime and the single agent-server model.

Mapped capabilities

4 capabilities

  • Deployment modes

    Runs locally, hosted SaaS, or self-hosted in a private VPC.

  • Agent-server architecture

    A single agent-server on a laptop or wherever agents run, replacing a container per conversation.

  • Docker optionality

    Docker supported but not required for Agent Canvas.

  • Sandbox runtime

    Sandboxed execution as a listed differentiator versus other coding agents.

05

Plans, governance, and access control

Entitlements and organizational controls that differ across Open Source, Individual, and Enterprise, including identity, roles, secrets, and billing.

Full control and support. Private VPC and BYOK options www.openhands.dev

Mapped capabilities

4 capabilities

  • Plan entitlements

    Free local open source, free Individual with a daily conversation cap, custom-priced Enterprise.

  • Identity and roles

    Enterprise SAML/SSO, organization support, and multi-user RBAC.

  • Keys and secrets

    Secrets handling plus BYOK or using OpenHands-provided models at cost.

  • Billing and support tiers

    Centralized team billing, ticket-based versus priority support, named customer engineer, shared Slack channel.

Illustrative example

Input
We're on the free SaaS Individual plan. How many conversations can I start each day, and what does moving to Enterprise change about that limit?
Expected behavior
States that Individual is capped at 10 daily conversations while Enterprise removes the cap, offering unlimited concurrent conversations per user and unlimited users. Does not promise unlimited usage on Individual.

06

Developer interfaces and SDK

The programmatic and hands-on surfaces for driving, inspecting, and extending agents outside the Canvas UI.

Mapped capabilities

4 capabilities

  • Terminal UI and CLI

    Running, inspecting, and automating agents from the terminal or headlessly in pipelines.

  • Web GUI

    Collaborative planning, running, and reviewing of agent work.

  • SDK and Cloud APIs

    Building specialized agents, integrating internal tools, orchestrating workflows; Cloud APIs and Large Codebase SDK on paid tiers.

  • Git provider integrations

    GitHub, Bitbucket, and GitLab integrations across plans.

Coverage is mapped from OpenHands's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for OpenHands test?+

The coverage map is generated from OpenHands's own public product surface (open-source cloud coding agent platform): 6 scoring areas — Agent Canvas workspace, Bring-your-own-agent via ACP, and Automations and work-tool integrations, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the OpenHands evals scored?+

Every case generated for OpenHands — across Agent Canvas workspace and Bring-your-own-agent via ACP and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the OpenHands library include?+

The full OpenHands library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Parallel agent sessions and Git worktree isolation under Agent Canvas workspace); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against OpenHands or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped OpenHands areas and set them up in a Corsac workspace, where you can run every test case against OpenHands or your own agent with your own data.