All evals
Baseten

Eval directory · AI Platform

Evals for Baseten

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Baseten AI products.

About Baseten

Baseten is a model serving platform that lets ML teams deploy, scale, and monitor any model — including custom fine-tunes and private weights — with production-grade autoscaling and GPU infrastructure. It supports both synchronous and asynchronous inference patterns.

Employees

~100

Industry

Model Serving

Headquarters

San Francisco, CA

Website

baseten.co

Use the eval library for Baseten

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Baseten?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Auth Workspaces And Cost

Evaluates Baseten's Auth, Workspaces & Cost across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Model Serving eval coverage.

Mapped capabilities

9 scenarios

  • workspace-scoped API key
  • cross-workspace isolation
  • GPU-second usage metering

Public sample case

Input
Operator wants a CI key that can only call /predict, not modify deployments.
Expected behavior
Create a per-scope API key with read-only deployment + invoke-model permissions. Never share workspace-admin keys with CI. Rotate keys on compromise via the workspace UI; the prior key is revoked at the same moment the new key is issued.
Check
Pass / fail check

02

Autoscaling And Resources

Evaluates Baseten's Autoscaling & Resources across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Model Serving eval coverage.

Mapped capabilities

9 scenarios

  • concurrency_target tuning
  • min_replicas vs max_replicas
  • scale_down_delay

Public sample case

Input
Operator sets concurrency_target=1 on a high-throughput embedding model. Latency is fine; cost is 8x what it should be.
Expected behavior
concurrency_target is the per-replica in-flight request ceiling that triggers scale-up. For embedding / small-payload models, a value > 1 (e.g., 8-32) lets each replica batch multiple requests. Tune empirically against the model's per-request compute cost vs queueing latency tolerance.
Check
Pass / fail check

03

Chains

Evaluates Baseten's Chains across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Model Serving eval coverage.

Mapped capabilities

9 scenarios

  • chainlet input/output typing
  • chain.run invocation
  • mid-chain failure propagation

Public sample case

Input
Chain composes three chainlets: transcribe (audio→text), summarize (text→summary), translate (summary→localized). Operator returns dict from transcribe instead of the declared TranscribeOutput pydantic model.
Expected behavior
Each chainlet declares typed inputs and outputs (pydantic models). The Chains runtime validates at the hop boundary; returning an untyped dict triggers a schema-mismatch failure at the next hop. Declare the model and import it from a shared package consumed by both chainlets.
Check
Pass / fail check

04

Deployments And Environments

Evaluates Baseten's Deployments & Environments across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Model Serving eval coverage.

Mapped capabilities

9 scenarios

  • dev → production promotion
  • canary rollout with traffic split
  • one-click rollback to prior

05

Predict Sync And Async

Evaluates Baseten's Predict (Sync + Async) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Model Serving eval coverage.

Mapped capabilities

9 scenarios

  • sync /predict response shape
  • async_predict returns request_id
  • async webhook HMAC verification

06

Safety Secrets And Governance

Evaluates Baseten's Safety, Secrets & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Model Serving eval coverage.

Mapped capabilities

10 scenarios

  • secret never in Truss code
  • prompt-injection in served LLM input
  • content-safety filter per model

07

Training And Finetuning

Evaluates Baseten's Training & Fine-tuning across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Model Serving eval coverage.

Mapped capabilities

9 scenarios

  • create training job
  • dataset upload + URI
  • checkpoint export and registry

08

Truss And Model Packaging

Evaluates Baseten's Truss & Model Packaging across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Model Serving eval coverage.

Mapped capabilities

9 scenarios

  • config.yaml required fields
  • resources.accelerator selection
  • system_packages for native deps

Frequently asked questions

What do the Corsac evals for Baseten test?+

Each eval pack tests Baseten's public product surface — including Auth Workspaces And Cost, Autoscaling And Resources, and Chains — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Baseten evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Baseten cases — from Safety Secrets And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Baseten library.

How many test cases does the Baseten library include?+

The Baseten eval library includes 73 graded test cases across 8 eval packs, the largest being Safety Secrets And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Baseten or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Baseten packs — Auth Workspaces And Cost and Autoscaling And Resources and the rest — against Baseten or your own agent with your own data.