All evals
Portkey

Eval directory · AI Platform

Evals for Portkey

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Portkey AI products.

About Portkey

Portkey is an AI gateway for production LLM apps — a unified, OpenAI-compatible API across 200+ models with provider routing and fallbacks, semantic and simple caching, input/output guardrails (PII redaction, prompt-injection, content moderation), request-level observability and traces, a versioned prompt library, virtual keys with per-key budgets and rate limits, and workspace RBAC + audit logs.

Employees

~40

Industry

AI Gateway

Headquarters

San Francisco, CA

Website

portkey.ai

Use the eval library for Portkey

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Portkey?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Caching Simple And Semantic

Evaluates Portkey's Simple & Semantic across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Gateway eval coverage.

Mapped capabilities

9 scenarios

  • cache.mode=simple exact match
  • cache.mode=semantic similarity_threshold
  • cache hit verification via headers

Public sample case

Input
Config cache={mode:'simple', max_age:3600}. Caller sends two requests with identical messages[] and model; one has temperature=0, the other temperature=0.1.
Expected behavior
Simple cache keys on the full request payload — different temperature values are different keys and miss. Verify with x-portkey-cache-status: MISS on both requests. To collapse them, normalize the request body client-side before sending.
Check
Pass / fail check

02

Configs And Virtual Keys

Evaluates Portkey's Configs & Virtual Keys across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Gateway eval coverage.

Mapped capabilities

10 scenarios

  • virtual key abstracts provider credentials
  • virtual key budget exceeded
  • virtual key rate limit per window

Public sample case

Input
Team has OpenAI, Anthropic, Bedrock credentials; wants to rotate the OpenAI key without code changes.
Expected behavior
Create a virtual key per upstream credential set; clients pass x-portkey-virtual-key (or it's set in a config) and never see the raw provider credential. Rotation = re-issue the underlying credential in the virtual key — operator code is unchanged.
Check
Pass / fail check

03

Guardrails Input And Output

Evaluates Portkey's Input & Output Guards across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Gateway eval coverage.

Mapped capabilities

9 scenarios

  • before_request_hooks ordering
  • PII redaction with sanitize action
  • prompt-injection guard deny action

Public sample case

Input
Config has before_request_hooks=[pii_redact, prompt_injection_detect]. Question: order of evaluation and short-circuit behavior.
Expected behavior
Hooks execute in declared order; first hook with action=deny short-circuits the request and the gateway returns the guard's denial response. Verify by sending a payload that fails the second hook only — observe it reaches the second hook because the first passed.
Check
Pass / fail check

04

Observability Logs And Traces

Evaluates Portkey's Observability, Logs & Traces across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Gateway eval coverage.

Mapped capabilities

9 scenarios

  • trace_id propagation across multi-step agent
  • x-portkey-metadata for per-tenant breakdown
  • latency p95 attribution by upstream

05

Prompt Library And Versioning

Evaluates Portkey's Prompt Library & Versioning across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Gateway eval coverage.

Mapped capabilities

9 scenarios

  • versioned prompt pin
  • variable substitution shape
  • render-only endpoint vs completion

06

Provider Routing And Fallbacks

Evaluates Portkey's Provider Routing & Fallbacks across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Gateway eval coverage.

Mapped capabilities

9 scenarios

  • strategy=fallback ordered targets
  • strategy=loadbalance with weights
  • strategy=conditional with query rules

07

Safety Rbac And Governance

Evaluates Portkey's Safety, RBAC & Governance across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Gateway eval coverage.

Mapped capabilities

9 scenarios

  • workspace member roles
  • audit log capture of admin actions
  • cross-workspace isolation

08

Unified Api And Gateway Proxy

Evaluates Portkey's Unified API & Gateway Proxy across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Gateway eval coverage.

Mapped capabilities

9 scenarios

  • OpenAI SDK base_url swap
  • x-portkey-api-key header required
  • model id namespacing across providers

Frequently asked questions

What do the Corsac evals for Portkey test?+

Each eval pack tests Portkey's public product surface — including Caching Simple And Semantic, Configs And Virtual Keys, and Guardrails Input And Output — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Portkey evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Portkey cases — from Configs And Virtual Keys (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Portkey library.

How many test cases does the Portkey library include?+

The Portkey eval library includes 73 graded test cases across 8 eval packs, the largest being Configs And Virtual Keys with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Portkey or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Portkey packs — Caching Simple And Semantic and Configs And Virtual Keys and the rest — against Portkey or your own agent with your own data.