01
Model APIs and OpenAI Compatibility
The serverless, pay-per-token API surface: drop-in OpenAI compatibility, structured generation features, and multi-modal request handling for production integration.
“JSON‑mode, function/tool calling, and schema‑guided outputs for consistent, structured results.” friendli.ai
Mapped capabilities
4 capabilities
Drop-in OpenAI client migration
Correctly explains that swapping the base URL lets existing OpenAI client code work, and what a caller still needs to change (API key, model id).
Structured output and tool calling
JSON mode, function/tool calling, and schema-guided outputs — what the feature does and how a request is shaped.
Multi-modality and agentic workflows
Text, vision, and other modalities served through a single API surface for agentic use.
Model catalog and availability on Model APIs
Which frontier open-weight models are pre-optimized and served serverless, without overstating availability beyond the supplied catalog.
Illustrative example
- Input
- Is FriendliAI actually 2x faster than other inference providers? We need to justify this to our VP.
- Expected behavior
- Reports 2x+ faster inference as FriendliAI's own marketing claim, attributed to its stack of custom kernels, caching, continuous batching, and speculative decoding. Notes it is not an independent measurement and recommends benchmarking on the buyer's own workload.






