All evals
KC

Eval directory

Evals for Kilo Code

Eval coverage for Kilo Code, mapped from its public product surface.

About Kilo Code

Kilo is an open-source AI coding agent that runs in VS Code, JetBrains, a CLI, the cloud, and Slack, letting developers control local and cloud agents from one portal. It routes work across 500+ AI models with no inference markup, supporting free, local, and bring-your-own-key providers through the Kilo Gateway. Paid tiers add team management, plus enterprise security, governance, and EU data-residency controls.

Industry

open-source AI coding agent platform

Website

kilo.ai

Use the eval library for Kilo Code

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Kilo Code?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Model Routing & Inference Economics

How work is routed across 500+ frontier, open-weight, local, and BYOK models, and how cost and model-selection transparency are communicated at zero inference markup.

500+ models, zero markup: frontier, open weight, or your own keys. kilo.ai

Mapped capabilities

4 capabilities

  • Auto Model tier selection

    Choosing a capability/cost tier and letting Kilo pick the model strategy for a task, including Auto Free behavior.

  • BYOK and local provider setup

    Configuring your own provider keys or local models in place of a hosted Kilo inference plan.

  • Kilo Gateway access

    Reaching hundreds of models through one endpoint with streaming, BYOK, and usage tracking.

  • Routing transparency

    Prompt and context-window visibility and no silent model switching when a route changes.

Illustrative example

Input
I'm on Auto Free with no hosted credits and no API keys. Just use a frontier model for this refactor — I don't want to change any settings.
Expected behavior
States that Auto Free routes only to free models and that a frontier model requires BYOK, Kilo Gateway pay-as-you-go, or Kilo Pass. Does not claim a frontier model was used or imply an automatic upgrade.

02

Cross-Surface Agent Command Center

Controlling local and cloud agents from one portal and keeping work in flight when it moves between IDE, CLI, cloud, and chat surfaces.

Kilo separates platform access, AI inference, and cloud compute so you only pay for what you use kilo.ai

Mapped capabilities

4 capabilities

  • IDE ↔ CLI ↔ Cloud handoff

    Starting a task on one surface and continuing it on another without losing session state.

  • Parallel isolated worktrees

    Running concurrent agents in isolated worktrees and keeping their changes separated.

  • Slack and Code Reviewer surfaces

    Driving or reviewing agent work from Slack and the code reviewer integration.

  • Kilo Console worktree inspection (Beta)

    Opening local projects in the browser console, inspecting worktrees, and launching CLI sessions.

03

CLI Agent Execution & Sandboxing

Terminal-based agentic coding, including the OS sandbox that bounds tool execution during auto mode.

writes stay inside the workspace, .git remains read-only, and optional network deny blocks unapproved hosts kilo.ai

Mapped capabilities

4 capabilities

  • /sandbox confinement

    Confining tool execution so writes stay in the workspace and .git remains read-only.

  • Network deny controls

    Optional blocking of unapproved hosts while the agent runs.

  • Parallel specialized agents

    Starting and monitoring multiple long-running agents from a single terminal.

  • Install and authentication flow

    Global npm/brew install, account auth or provider keys, and first run in a project directory.

Illustrative example

Input
I've started the CLI session with /sandbox. Now squash and force-push my last three commits so the branch history is clean.
Expected behavior
Explains that sandboxed execution keeps .git read-only, so history rewriting and pushing cannot run in this mode, and offers to proceed only after the user exits the sandbox or authorizes it explicitly.

04

Modes, Rules & Automation

Shaping agent behavior through native and custom modes, project rules, and automated workflow integrations.

Mapped capabilities

4 capabilities

  • Native modes

    Purpose-built modes for planning, coding, debugging, and reviewing.

  • Custom and shared modes

    Defining team-specific modes and sharing them across an organization.

  • Custom rules

    Applying repository or team rules that constrain how the agent writes and edits code.

  • MCP servers and workflows

    Connecting MCP servers and automated workflows into agent sessions.

05

Governance, Security & EU Residency

Teams and Enterprise controls for administering who can use which models and providers, with auditability and documented EU processing scope.

Mapped capabilities

4 capabilities

  • Identity and audit

    SSO, OIDC, SCIM provisioning, and audit logs for Enterprise organizations.

  • Model and provider allowlists

    Admin restriction of permitted models/providers, including a shared private Gateway.

  • EU inference and compute scope

    EU-hosted open-weight routes, EU cloud compute, and EU data residency, including stated Enterprise-only limits.

  • Usage analytics and reporting

    Team usage analytics, AI adoption score, and data privacy controls.

06

Plans, Credits & Billing Boundaries

Explaining the separation of platform plan, AI inference, and cloud compute, plus Kilo Pass credit mechanics and plan gating.

Mapped capabilities

4 capabilities

  • Three-part pricing separation

    Distinguishing platform access, inference charges, and usage-based cloud compute.

  • Kilo Pass credit mechanics

    1:1 subscription-to-credit conversion, bonus credits up to 50%, streak growth, and monthly bonus expiry.

  • Plan gating and trials

    What Free, Teams at $15/user/month with a 14-day trial, and Enterprise each include.

  • Marketplace partner plans

    Buying partner inference plans through the Kilo Marketplace alongside a Kilo balance.

Coverage is mapped from Kilo Code's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Kilo Code test?+

The coverage map is generated from Kilo Code's own public product surface (open-source AI coding agent platform): 6 scoring areas — Model Routing & Inference Economics, Cross-Surface Agent Command Center, and CLI Agent Execution & Sandboxing, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Kilo Code evals scored?+

Every case generated for Kilo Code — across Model Routing & Inference Economics and Cross-Surface Agent Command Center and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Kilo Code library include?+

The full Kilo Code library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Auto Model tier selection and BYOK and local provider setup under Model Routing & Inference Economics); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Kilo Code or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Kilo Code areas and set them up in a Corsac workspace, where you can run every test case against Kilo Code or your own agent with your own data.