All evals
Fireworks AI

Eval directory · AI Platform

Evals for Fireworks AI

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Fireworks AI AI products.

About Fireworks AI

Fireworks AI is a high-performance inference platform for open-source and fine-tuned models, delivering industry-leading throughput and latency for production workloads. Teams use Fireworks to run Llama, Mixtral, and custom fine-tunes at scale without managing GPU infrastructure.

Employees

~80

Industry

AI Inference

Headquarters

San Francisco, CA

Use the eval library for Fireworks AI

All 74 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Fireworks AI?

6 areas · 74 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Fireworks Batch Prompt Cache Runtime Performance

Evaluates Fireworks AI's Batch, Prompt Cache & Runtime Performance across 13 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI infrastructure eval coverage.

Mapped capabilities

13 scenarios

  • Prompt cache prefix
  • Prompt cache discipline
  • Streaming TTFT

Public sample case

Input
Same 8k-token document prefix across requests; cache should reduce cost on shared prefix.
Expected behavior
Place static RAG context in stable system message prefix; keep variable user query suffix; rely on documented prompt cache behavior.
Check
Pass / fail check

02

Fireworks Deployment Topology Capacity

Evaluates Fireworks AI's Deployment Topology & Capacity across 13 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI infrastructure eval coverage.

Mapped capabilities

13 scenarios

  • Serverless vs on-demand
  • GPU SKU sizing
  • Region and GPU SKU selection

Public sample case

Input
Prototype uses shared serverless; production SLA needs predictable capacity via firectl deployment create.
Expected behavior
Recommend on-demand or dedicated deployment via firectl when SLA requires reserved capacity; keep serverless for bursty dev traffic.
Check
Pass / fail check

03

Fireworks Fine Tuning Multi Lora Serving

Evaluates Fireworks AI's Fine-Tuning & Multi-LoRA Serving across 12 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI infrastructure eval coverage.

Mapped capabilities

12 scenarios

  • FireOptimizer tuning
  • Adapter version pinning
  • Multiple adapters per deployment

Public sample case

Input
Legal domain fine-tune needs conservative learning rate; agent configures job not inference API.
Expected behavior
Set FireOptimizer fine-tuning job parameters per docs; evaluate adapter on holdout before Multi-LoRA deploy.
Check
Pass / fail check

04

Fireworks Function Calling Tool Orchestration

Evaluates Fireworks AI's Function Calling & Tool Orchestration across 12 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI infrastructure eval coverage.

Mapped capabilities

12 scenarios

  • Multi-step agent loop
  • tool_choice control
  • Invalid tool arguments

05

Fireworks Safety Moderation Observability

Evaluates Fireworks AI's Safety, Moderation & Observability across 11 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI infrastructure eval coverage.

Mapped capabilities

11 scenarios

  • OTel mapping
  • Usage fields
  • Latency observability

06

Fireworks Structured Outputs Grammar Constraints

Evaluates Fireworks AI's Structured Outputs & Grammar Constraints across 13 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI infrastructure eval coverage.

Mapped capabilities

13 scenarios

  • JSON schema mode

Frequently asked questions

What do the Corsac evals for Fireworks AI test?+

Each eval pack tests Fireworks AI's public product surface — including Fireworks Batch Prompt Cache Runtime Performance, Fireworks Deployment Topology Capacity, and Fireworks Fine Tuning Multi Lora Serving — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Fireworks AI evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 74 Fireworks AI cases — from Fireworks Batch Prompt Cache Runtime Performance (13 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Fireworks AI library.

How many test cases does the Fireworks AI library include?+

The Fireworks AI eval library includes 74 graded test cases across 6 eval packs, the largest being Fireworks Batch Prompt Cache Runtime Performance with 13 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Fireworks AI or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 6 Fireworks AI packs — Fireworks Batch Prompt Cache Runtime Performance and Fireworks Deployment Topology Capacity and the rest — against Fireworks AI or your own agent with your own data.