
Byok Isolation Usage Provenance
OpenRouter · OpenRouter
LLM routing and aggregation — OpenRouter
Evaluates OpenRouter's BYOK Isolation & Usage Provenance across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM routing and aggregation eval coverage.
About OpenRouter
OpenRouter is a unified LLM routing layer that gives developers access to hundreds of models through a single OpenAI-compatible API. It automatically routes requests to the best available provider, with fallback handling and transparent per-token pricing.
Sample tests· showing 3 of 9
| # | Input | Expected behavior | Check |
|---|---|---|---|
| 01 | Tenant uploaded Anthropic key; response usage shows is_byok true and cost near zero on OpenRouter ledger while upstream_inference_cost populated in cost_details. | Map is_byok true to pass-through upstream billing for tenant key spend while separating OpenRouter platform fees per documented cost_details fields. | Pass / FailBillingcritical |
| 02 | Dashboard shows active Anthropic BYOK; application still sends OPENROUTER_API_KEY only; expects tenant key consumption. | OpenRouter prioritizes tenant BYOK for matching provider before platform pooled keys, reflected in is_byok and provenance metadata when enabled. | Pass / FailRoutinghigh |
| 03 | MSP hosts multiple customers under one org; each customer workspace registers separate OpenAI BYOK keys. | Keys are workspace-scoped; routing and usage attribution never bleed across workspaces even within shared org billing. | Pass / FailSecuritycritical |
How this eval is graded
Grade against expected.ideal_behavior and expected.rubric.
Rubric criteria
- Openrouter
- Ai Platform
- Byok Isolation Usage Provenance
Recommended for
Works with
Related evals
Claude API
Evaluates Anthropic's Batch API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.
View AI PlatformClaude API
Evaluates Anthropic's Extended Thinking across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.
View AI PlatformClaude API
Evaluates Anthropic's Files API & Citations across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.
ViewFrequently asked questions
What does the Byok Isolation Usage Provenance eval for OpenRouter OpenRouter test?+
Evaluates OpenRouter's BYOK Isolation & Usage Provenance across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM routing and aggregation eval coverage.
How is the Byok Isolation Usage Provenance eval scored?+
The judge rubric: Grade against expected.ideal_behavior and expected.rubric.
How many test cases does this eval pack include?+
The Byok Isolation Usage Provenance pack for OpenRouter OpenRouter contains 9 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.
How do I run this eval?+
Sign up for Corsac, connect your model or agent endpoint, and run the Byok Isolation Usage Provenance pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.
Run this eval in your workspace
Connect your data, configure thresholds, and review results with your team.