Eval directory
Evals for fal
8 evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for fal AI products.
About fal
fal is a generative media inference platform offering fast, scalable image, video, and audio generation through a simple API. It hosts leading open models (FLUX, Stable Diffusion, Whisper) and supports fine-tuned LoRA routing, webhooks, and queue-based async generation at production scale.
Available eval packs for fal
8 packs ready to run.
Billing Concurrency Media Lifecycle
fal evals — Billing Concurrency Media Lifecycle (relift v3)
Content Safety Refusal Rights
fal evals — Content Safety Refusal Rights (relift v3)
Image Video Audio Generation
fal evals — Image Video Audio Generation (relift v3)
Lora Routing Private Models
fal evals — Lora Routing Private Models (relift v3)
Model Catalog Seed Reproducibility
fal evals — Model Catalog Seed Reproducibility (relift v3)
Queue Webhooks Reliability
fal evals — Queue Webhooks Reliability (relift v3)
Serverless Deploy Tenant Isolation
fal evals — Serverless Deploy Tenant Isolation (relift v3)
Workflows Chained Pipelines
fal evals — Workflows Chained Pipelines (relift v3)
Why eval fal AI
fal's AI features ship behind brand promises about accuracy, safety, and reliability. Buyers and integrators need to know those promises hold up under adversarial prompts, edge-case workflows, and the long tail of real customer inputs — not just the demo path.
The Corsac eval library for fal measures four dimensions teams care about most when deploying ai platform agents:
- Adversarial robustness — does the agent resist prompt injection, jailbreaks, and social-engineering attempts?
- Workflow quality— does it complete the task buyers were shown in the demo, on inputs that don't look like the demo?
- Safety gates — does it escalate or refuse when it should, and only then?
- Operator quality — does it preserve analyst trust by surfacing the right context at the right time?
Every eval pack above is hand-authored against fal's public product surface and runnable in Corsac with your own data.