All evals
OpenRouter

Eval directory · AI Platform

Evals for OpenRouter

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for OpenRouter AI products.

About OpenRouter

OpenRouter is a unified LLM routing layer that gives developers access to hundreds of models through a single OpenAI-compatible API. It automatically routes requests to the best available provider, with fallback handling and transparent per-token pricing.

Employees

~20

Industry

LLM Infrastructure

Headquarters

United States

Use the eval library for OpenRouter

All 68 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for OpenRouter?

7 areas · 68 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Byok Isolation Usage Provenance

Evaluates OpenRouter's BYOK Isolation & Usage Provenance across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM routing and aggregation eval coverage.

Mapped capabilities

9 scenarios

  • usage.is_byok flag
  • workspace key precedence
  • cross-tenant key isolation

Public sample case

Input
Tenant uploaded Anthropic key; response usage shows is_byok true and cost near zero on OpenRouter ledger while upstream_inference_cost populated in cost_details.
Expected behavior
Map is_byok true to pass-through upstream billing for tenant key spend while separating OpenRouter platform fees per documented cost_details fields.
Check
Pass / fail check

02

Model Catalog Alias Stability

Evaluates OpenRouter's Model Catalog & Alias Stability across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM routing and aggregation eval coverage.

Mapped capabilities

10 scenarios

  • GET /models freshness
  • :floor price variant
  • ~latest alias drift

Public sample case

Input
Cron caches models JSON for routing agent; new provider added mid-day for meta-llama slug.
Expected behavior
Recommend periodic refresh with versioned ETag or timestamped cache invalidation; stale cache risks require_parameters mismatch errors at runtime.
Check
Pass / fail check

03

Parameter Parity Capability Routing

Evaluates OpenRouter's Parameter Parity & Capability Routing across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM routing and aggregation eval coverage.

Mapped capabilities

10 scenarios

  • tools provider filter
  • json_schema response_format
  • parallel_tool_calls flag

Public sample case

Input
Five providers listed for openai/gpt-4o; three lack tools in supported_parameters; tools array populated.
Expected behavior
Router restricts to tool-capable providers automatically; if none after only/ignore filters, fail with explicit error.
Check
Pass / fail check

04

Privacy Data Policy Routing

Evaluates OpenRouter's Privacy & Data-Policy Routing across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM routing and aggregation eval coverage.

Mapped capabilities

10 scenarios

  • provider.zdr routing
  • data_collection deny
  • enforce_distillable_text

05

Provider Fallback Outage Routing

Evaluates OpenRouter's Provider Fallback & Outage Routing across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM routing and aggregation eval coverage.

Mapped capabilities

10 scenarios

  • 5xx provider failover
  • allow_fallbacks false
  • models[] fallback chain

06

Router Metadata Observability

Evaluates OpenRouter's Router Metadata & Observability across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM routing and aggregation eval coverage.

Mapped capabilities

10 scenarios

  • experimental metadata header
  • pipeline attempts array
  • cache metadata strip

07

Usage Accounting Cost Attribution

Evaluates OpenRouter's Usage Accounting & Cost Attribution across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM routing and aggregation eval coverage.

Mapped capabilities

9 scenarios

  • usage.cost billing
  • cached_tokens discount
  • generation audit API

Frequently asked questions

What do the Corsac evals for OpenRouter test?+

Each eval pack tests OpenRouter's public product surface — including Byok Isolation Usage Provenance, Model Catalog Alias Stability, and Parameter Parity Capability Routing — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the OpenRouter evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 68 OpenRouter cases — from Model Catalog Alias Stability (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the OpenRouter library.

How many test cases does the OpenRouter library include?+

The OpenRouter eval library includes 68 graded test cases across 7 eval packs, the largest being Model Catalog Alias Stability with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against OpenRouter or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 7 OpenRouter packs — Byok Isolation Usage Provenance and Model Catalog Alias Stability and the rest — against OpenRouter or your own agent with your own data.