All evals
Perplexity

Eval directory · AI Platform

Evals for Perplexity

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Perplexity AI products.

About Perplexity

Perplexity is an answer engine; the Perplexity Sonar API exposes its grounded LLM with real-time web search and inline citations — sonar, sonar-pro, and sonar-reasoning models, source filtering and recency controls, and OpenAI-compatible chat completions for grounded answers at API scale.

Employees

~200

Industry

Search / Answer API

Headquarters

San Francisco, CA

Use the eval library for Perplexity

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Perplexity?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Auth Rate Limits Tiers And Cost

Evaluates Perplexity's Auth, Rate Limits, Tiers & Cost across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Grounded Answer API eval coverage.

Mapped capabilities

10 scenarios

  • Bearer token in Authorization header
  • key rotation without downtime
  • 429 + Retry-After backoff

Public sample case

Input
Operator sets Authorization: 'Bearer PPLX_...' on every request to api.perplexity.ai.
Expected behavior
Use Authorization: Bearer <PPLX_API_KEY>. The key is the Perplexity-issued one (not OpenAI's). Store in secret manager, rotate on suspected exposure, never log raw key. Verify the header is sent on every request (no silent unauthenticated fallback).
Check
Pass / fail check

02

Chat Completions Openai Compatible

Evaluates Perplexity's Chat Completions (OpenAI-compatible) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Grounded Answer API eval coverage.

Mapped capabilities

9 scenarios

  • model selection sonar vs sonar-pro vs sonar-reasoning
  • OpenAI SDK base_url override
  • streaming SSE OpenAI delta format

Public sample case

Input
Operator hardcodes model='sonar' for every workload — quick factual lookups, deep multi-hop research, and structured-output extractions — to 'keep things simple.'
Expected behavior
Route by workload: sonar for low-latency single-hop factual answers; sonar-pro for higher-quality multi-source synthesis with deeper grounding; sonar-reasoning for chain-of-thought heavy tasks where latency budget allows. Encode the choice per route, do not hardcode.
Check
Pass / fail check

03

Citations And Source Grounding

Evaluates Perplexity's Citations & Source Grounding across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Grounded Answer API eval coverage.

Mapped capabilities

9 scenarios

  • return_citations=true enables citations[]
  • in-text [n] anchor → citations[i] mapping
  • citation faithfulness check (claim vs source)

Public sample case

Input
Operator omits return_citations on POST /chat/completions and expects the response.citations[] array to populate by default.
Expected behavior
Set return_citations=true explicitly on every request where source attribution is required downstream. The default behavior should not be assumed [REQUIRES-VERIFICATION] across model versions — opt in by setting the flag and verify response.citations is present and non-empty before rendering.
Check
Pass / fail check

04

Images And Multimodal

Evaluates Perplexity's Images & Multimodal across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Grounded Answer API eval coverage.

Mapped capabilities

9 scenarios

  • return_images=true appends images[]
  • image URL host vs search_domain_filter scope
  • image input via messages content parts (vision)

05

Reasoning Models Sonar Reasoning

Evaluates Perplexity's Reasoning Models (Sonar Reasoning) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Grounded Answer API eval coverage.

Mapped capabilities

9 scenarios

  • sonar-reasoning latency budget
  • reasoning tokens count as output
  • reasoning content vs final answer extraction

06

Safety Policy Source Quality And Governance

Evaluates Perplexity's Safety, Policy, Source Quality & Governance across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Grounded Answer API eval coverage.

Mapped capabilities

9 scenarios

  • AUP-prohibited request handling
  • prompt injection via retrieved page
  • source-quality discipline (citation hosts)

07

Search Controls

Evaluates Perplexity's Search Controls across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Grounded Answer API eval coverage.

Mapped capabilities

9 scenarios

  • search_domain_filter allow-list
  • search_domain_filter deny-list (- prefix)
  • search_recency_filter='day' for breaking news

08

Structured Outputs And Format Controls

Evaluates Perplexity's Structured Outputs & Format Controls across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Grounded Answer API eval coverage.

Mapped capabilities

9 scenarios

  • response_format json_schema basic shape
  • regex response_format
  • schema under grounded-context conflict

Frequently asked questions

What do the Corsac evals for Perplexity test?+

Each eval pack tests Perplexity's public product surface — including Auth Rate Limits Tiers And Cost, Chat Completions Openai Compatible, and Citations And Source Grounding — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Perplexity evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Perplexity cases — from Auth Rate Limits Tiers And Cost (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Perplexity library.

How many test cases does the Perplexity library include?+

The Perplexity eval library includes 73 graded test cases across 8 eval packs, the largest being Auth Rate Limits Tiers And Cost with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Perplexity or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Perplexity packs — Auth Rate Limits Tiers And Cost and Chat Completions Openai Compatible and the rest — against Perplexity or your own agent with your own data.