All evals
FriendliAI

Eval directory · AI Platform

Evals for FriendliAI

Eval coverage for FriendliAI, mapped from its public product surface.

About FriendliAI

FriendliAI is an inference platform for deploying frontier open-weight and custom AI models, offered as serverless Model APIs, Dedicated Endpoints on rented GPUs, and a self-hosted Friendli Container. It markets a purpose-built inference stack using custom GPU kernels, caching, continuous batching, and speculative decoding for lower latency and cost. Pricing is self-serve per-token or per-second GPU billing (A100 through B300), with a customizable Enterprise framework for VPC, on-prem, and reserved capacity.

Industry

AI inference cloud / LLM serving infrastructure

Use the eval library for FriendliAI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for FriendliAI?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Model APIs and OpenAI Compatibility

The serverless, pay-per-token API surface: drop-in OpenAI compatibility, structured generation features, and multi-modal request handling for production integration.

JSON‑mode, function/tool calling, and schema‑guided outputs for consistent, structured results. friendli.ai

Mapped capabilities

4 capabilities

  • Drop-in OpenAI client migration

    Correctly explains that swapping the base URL lets existing OpenAI client code work, and what a caller still needs to change (API key, model id).

  • Structured output and tool calling

    JSON mode, function/tool calling, and schema-guided outputs — what the feature does and how a request is shaped.

  • Multi-modality and agentic workflows

    Text, vision, and other modalities served through a single API surface for agentic use.

  • Model catalog and availability on Model APIs

    Which frontier open-weight models are pre-optimized and served serverless, without overstating availability beyond the supplied catalog.

Illustrative example

Input
Is FriendliAI actually 2x faster than other inference providers? We need to justify this to our VP.
Expected behavior
Reports 2x+ faster inference as FriendliAI's own marketing claim, attributed to its stack of custom kernels, caching, continuous batching, and speculative decoding. Notes it is not an independent measurement and recommends benchmarking on the buyer's own workload.

02

Model Deployment and Custom Models

Getting a specific model running: one-click deployment from the 600,000-model Hugging Face catalog, and bringing fine-tuned or proprietary weights.

Instantly deploy any of 600,000 Hugging Face models — from language to audio to vision — with a single click. friendli.ai

Mapped capabilities

3 capabilities

  • Hugging Face one-click deployment

    Deploying any of the 600K+ catalog models across language, audio, and vision without manual setup or optimization.

  • Bring-your-own fine-tuned or proprietary model

    Path for custom weights onto Dedicated Endpoints or Container, including where support is hands-on rather than self-serve.

  • Automatic optimization and tuning

    What the platform handles for the user (deployment, scaling, performance tuning) versus what the user configures.

03

Deployment Mode Selection

Choosing among serverless Model APIs, on-demand dedicated GPUs, reserved enterprise capacity, and the self-hosted container — the central architectural decision a buyer makes.

Mapped capabilities

4 capabilities

  • Serverless vs. dedicated crossover

    When workload scale justifies moving from per-token serverless to dedicated GPU capacity, framed as a workload-dependent decision.

  • On-demand vs. reserved capacity

    Per-second on-demand GPU billing versus 1+ month reserved instances with discounted upfront payment and exclusive features.

  • Container vs. managed cloud

    Self-hosted Friendli Container for data residency and infrastructure control versus FriendliAI-managed endpoints.

  • Migration path between modes

    Moving from Model APIs to Dedicated Endpoints for predictable throughput and isolation.

04

Pricing and Billing Mechanics

Self-serve per-token and per-second GPU pricing, the published GPU rate card, and the boundary where Enterprise becomes a contract conversation.

Only pay for the compute you use, down to the second, with no extra charges for start-up times friendli.ai

Mapped capabilities

4 capabilities

  • GPU rate card accuracy

    A100 80GB $2.9/hr, H100 80GB $3.9/hr, H200 141GB $4.5/hr, B200 180GB $8.9/hr, B300 288GB $12.0/hr, billed per second.

  • Per-second billing and start-up time

    On-demand deployments pay only for compute used, down to the second, with no extra charge for start-up time.

  • Cost estimation for a stated workload

    Arithmetic on published rates for a given GPU type and duration, without inventing token prices not present in the context.

  • Self-serve vs. contact-sales boundary

    Model APIs and Dedicated Endpoints are self-serve; Container pricing and Enterprise terms require contacting sales.

Illustrative example

Input
We want one H200 141GB on-demand for 30 hours this month, plus a Friendli Container for our on-prem cluster. What will that cost?
Expected behavior
Computes the H200 portion as 30 x $4.5 = $135, noting per-second billing with no start-up charge. States that Container pricing is not published and requires contacting FriendliAI, rather than estimating it.

05

Reliability, Scale, and Performance Claims

How the platform's uptime, geo-distribution, failover, and speed claims are represented — including attribution and the limits of single-benchmark comparisons.

Ensure 99.99% uptime with our geo-distributed, multi-cloud infrastructure, engineered for reliability at scale. friendli.ai

Mapped capabilities

4 capabilities

  • Uptime SLA and failover architecture

    99.99% uptime SLA, geo-distributed multi-cloud infrastructure, active redundancy, and automated failover as marketed commitments.

  • Inference optimization techniques

    Custom GPU kernels, smart caching, continuous batching, speculative decoding, and parallel inference — described without inflating measured effect.

  • Attributing performance claims

    Presenting '2x+ faster', '#1 throughput on OpenRouter', and customer-reported savings as vendor or customer claims rather than independent measurement.

  • Traffic spikes and autoscaling

    Behavior under unpredictable traffic spikes and scaling across regions and GPU fleets.

06

Enterprise Control and Self-Hosted Deployment

The enterprise framework and Friendli Container: data residency, private networking, and commercial commitments — enabled by contract rather than as a fixed bundle.

Mapped capabilities

4 capabilities

  • VPC, on-prem, and custom regions

    Deployment topologies where data and models stay inside customer infrastructure.

  • Enterprise framework, not a fixed plan

    Features and capabilities are enabled per contract; avoid presenting Enterprise as a preset feature list.

  • Priority GPU access and custom rate limits

    Custom Model APIs rate limits, priority access to high-demand GPU types, and reserved capacity.

  • Support and commercial commitments

    Dedicated support channels, named Customer Success ownership, and custom commercial terms.

Coverage is mapped from FriendliAI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for FriendliAI test?+

The coverage map is generated from FriendliAI's own public product surface (AI inference cloud / LLM serving infrastructure): 6 scoring areas — Model APIs and OpenAI Compatibility, Model Deployment and Custom Models, and Deployment Mode Selection, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the FriendliAI evals scored?+

Every case generated for FriendliAI — across Model APIs and OpenAI Compatibility and Model Deployment and Custom Models and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the FriendliAI library include?+

The full FriendliAI library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Drop-in OpenAI client migration and Structured output and tool calling under Model APIs and OpenAI Compatibility); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against FriendliAI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped FriendliAI areas and set them up in a Corsac workspace, where you can run every test case against FriendliAI or your own agent with your own data.