All evals
Replicate

Eval directory · AI Platform

Evals for Replicate

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Replicate AI products.

About Replicate

Replicate is an AI model-hosting platform — run thousands of community and custom Cog-packaged models (FLUX, SDXL, Llama, Whisper, custom fine-tunes) via a simple HTTP API with predictions, webhooks, streaming, deployments, and per-second billing.

Employees

~80

Industry

AI Inference Platform

Headquarters

San Francisco, CA

Use the eval library for Replicate

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Replicate?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Auth Billing Safety And Governance

Evaluates Replicate's Auth, Billing, Safety & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Model Hosting eval coverage.

Mapped capabilities

10 scenarios

  • Bearer token header shape
  • per-second billing meter
  • NSFW / CSAM AUP gate

Public sample case

Input
Integrator sends Authorization: Bearer <REPLICATE_API_TOKEN> and gets 401.
Expected behavior
Replicate's documented header form is Authorization: Token <REPLICATE_API_TOKEN> (or Bearer, depending on the docs revision; [REQUIRES-VERIFICATION] against the current reference). Use the form the SDK uses. Never log the header value. Surface 401 as 'check token scope and rotation', not as transie…
Check
Pass / fail check

02

Cog And Custom Model Push

Evaluates Replicate's Cog & Custom Model Push across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Model Hosting eval coverage.

Mapped capabilities

9 scenarios

  • cog.yaml gpu / hardware selection
  • predict() input typing
  • setup() vs predict() warm vs cold

Public sample case

Input
Integrator's cog.yaml declares build.gpu=true but does not specify a GPU class. Push succeeds but predictions OOM on the assigned tier.
Expected behavior
Match the predict.py memory footprint to the documented per-tier VRAM (T4 16GB, A40 48GB, A100 40/80GB, H100 80GB). Declare the target hardware tier on the model in the Replicate UI (or via the API) — build.gpu=true is necessary but not sufficient. Test with a representative input before shipping.
Check
Pass / fail check

03

Deployments

Evaluates Replicate's Deployments across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Model Hosting eval coverage.

Mapped capabilities

9 scenarios

  • min_replicas controls cold start
  • max_replicas autoscaling cap
  • deployment hardware selection

Public sample case

Input
Operator runs a customer-facing FLUX deployment with min_replicas=0 to save cost. First request after 10 minutes idle takes 30 s instead of 2 s.
Expected behavior
min_replicas=0 enables scale-to-zero, trading cold-start latency for idle cost. For low-latency UX, set min_replicas>=1 during business hours (scheduled) or accept the cold-start budget. Per-tier cold-start latency [REQUIRES-VERIFICATION] — measure on the chosen hardware.
Check
Pass / fail check

04

Fine Tuning Replicate Train

Evaluates Replicate's Fine-tuning (Replicate Train) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Model Hosting eval coverage.

Mapped capabilities

9 scenarios

  • trainings.create with destination
  • training input schema per base model
  • training data hosting

05

Models Versions And Schema

Evaluates Replicate's Models, Versions & Schema across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Model Hosting eval coverage.

Mapped capabilities

9 scenarios

  • model slug owner/name parsing
  • immutable version id pinning
  • OpenAPI input schema introspection

06

Predictions Api

Evaluates Replicate's Predictions API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Model Hosting eval coverage.

Mapped capabilities

9 scenarios

  • version vs model:version invocation
  • Prefer: wait sync mode
  • async polling backoff

07

Streaming Predictions

Evaluates Replicate's Streaming Predictions across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Model Hosting eval coverage.

Mapped capabilities

9 scenarios

  • urls.stream presence gate
  • SSE event types: output / logs / done
  • client disconnect billing

08

Webhooks

Evaluates Replicate's Webhooks across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Model Hosting eval coverage.

Mapped capabilities

9 scenarios

  • webhook_events_filter scope
  • signature verification HMAC-SHA256
  • idempotency by webhook-id

Frequently asked questions

What do the Corsac evals for Replicate test?+

Each eval pack tests Replicate's public product surface — including Auth Billing Safety And Governance, Cog And Custom Model Push, and Deployments — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Replicate evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Replicate cases — from Auth Billing Safety And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Replicate library.

How many test cases does the Replicate library include?+

The Replicate eval library includes 73 graded test cases across 8 eval packs, the largest being Auth Billing Safety And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Replicate or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Replicate packs — Auth Billing Safety And Governance and Cog And Custom Model Push and the rest — against Replicate or your own agent with your own data.