
Workflows Chained Pipelines
fal · fal
Generative media inference — fal
Evaluates fal's Workflows Chained Pipelines across 6 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Generative media inference eval coverage.
About fal
fal is a generative media inference platform offering fast, scalable image, video, and audio generation through a simple API. It hosts leading open models (FLUX, Stable Diffusion, Whisper) and supports fine-tuned LoRA routing, webhooks, and queue-based async generation at production scale.
Sample tests· showing 3 of 6
| # | Input | Expected behavior | Check |
|---|---|---|---|
| 01 | Upscale after generation workflow published. | Call workflow id; pass inputs matching stage schema; verify stage outputs chained. | Pass / FailWorkflowhigh |
| 02 | Automation sends blank string. | Validate locally; return 400 client-side without queue submit. | Pass / FailWorkflowmedium |
| 03 | Partial pipeline charged incorrectly. | Detect stage error; do not bill user for final deliverable; surface partial failure code. | Pass / FailWorkflowhigh |
How this eval is graded
Grade against expected.ideal_behavior and expected.rubric.
Rubric criteria
- Fal
- Ai Platform
- Workflows Chained Pipelines
Recommended for
Works with
Related evals
Claude API
Evaluates Anthropic's Batch API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.
View AI PlatformClaude API
Evaluates Anthropic's Extended Thinking across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.
View AI PlatformClaude API
Evaluates Anthropic's Files API & Citations across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.
ViewFrequently asked questions
What does the Workflows Chained Pipelines eval for fal fal test?+
Evaluates fal's Workflows Chained Pipelines across 6 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Generative media inference eval coverage.
How is the Workflows Chained Pipelines eval scored?+
The judge rubric: Grade against expected.ideal_behavior and expected.rubric.
How many test cases does this eval pack include?+
The Workflows Chained Pipelines pack for fal fal contains 6 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.
How do I run this eval?+
Sign up for Corsac, connect your model or agent endpoint, and run the Workflows Chained Pipelines pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.
Run this eval in your workspace
Connect your data, configure thresholds, and review results with your team.