All evals
Groq

Eval directory · AI Platform

Evals for Groq

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Groq AI products.

About Groq

Groq builds the LPU (Language Processing Unit) inference engine and GroqCloud — an OpenAI-compatible API that serves leading open models (Llama, Mixtral, Gemma, Qwen) at very high tokens-per-second with low, deterministic latency. Developers use GroqCloud for real-time chat, tool use, structured outputs, and speech-to-text without managing GPU infrastructure.

Employees

~300

Industry

AI Inference Platform

Headquarters

Mountain View, CA

Website

groq.com

Use the eval library for Groq

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Groq?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Auth Rate Limits And Tiers

Evaluates Groq's Auth, Rate Limits & Tiers across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Fast Inference eval coverage.

Mapped capabilities

9 scenarios

  • Bearer API key handling
  • 429 with retry-after
  • ratelimit headers for pacing

Public sample case

Input
The agent passes the GroqCloud key as an x-api-key header (copied from a different vendor) and gets a 401.
Expected behavior
Authenticate with Authorization: Bearer <GROQ_API_KEY>. Load the key from a secret store / environment variable, never hardcode it, and confirm the header form against docs — GroqCloud uses Bearer auth, not an x-api-key header.
Check
Pass / fail check

02

Batch Api

Evaluates Groq's Batch API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Fast Inference eval coverage.

Mapped capabilities

9 scenarios

  • JSONL input file shape
  • custom_id result mapping
  • completion window selection

Public sample case

Input
Agent submits a batch with a plain JSON array instead of newline-delimited JSON, one request per line.
Expected behavior
Build the batch input as JSONL — one request object per line, each with a unique custom_id, method, url (the target endpoint), and body. Upload it via the Files API and reference the file id when creating the batch. A JSON array is not valid JSONL.
Check
Pass / fail check

03

Chat Completions Openai Compatible

Evaluates Groq's Chat Completions (OpenAI-compatible) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Fast Inference eval coverage.

Mapped capabilities

9 scenarios

  • base_url points at GroqCloud
  • model id selection across families
  • messages[] role ordering

Public sample case

Input
Operator reuses an existing OpenAI SDK client and forgets to override base_url, so chat completion requests with a Groq API key go to api.openai.com instead of GroqCloud.
Expected behavior
Point the OpenAI-compatible client at base_url https://api.groq.com/openai/v1 and pass the GROQ_API_KEY as the Bearer key. Verify the configured base_url before the first call — a Groq key against the OpenAI host returns 401 and never reaches the LPU.
Check
Pass / fail check

04

Function Calling And Tool Use

Evaluates Groq's Function Calling & Tool Use across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Fast Inference eval coverage.

Mapped capabilities

9 scenarios

  • tools[] schema definition
  • tool_choice=auto vs required vs named
  • tool_call_id pairing

05

Safety Models And Governance

Evaluates Groq's Safety, Models & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Fast Inference eval coverage.

Mapped capabilities

10 scenarios

  • model deprecation lifecycle
  • Llama Guard content moderation
  • moderation verdict enforcement

06

Speech Whisper Stt

Evaluates Groq's Speech (Whisper STT) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Fast Inference eval coverage.

Mapped capabilities

9 scenarios

  • transcriptions endpoint and model
  • transcription vs translation
  • language hint parameter

07

Speed Streaming And Latency

Evaluates Groq's Speed, Streaming & Latency across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Fast Inference eval coverage.

Mapped capabilities

9 scenarios

  • SSE delta chunk accumulation
  • [DONE] sentinel handling
  • finish_reason on final chunk

08

Structured Outputs And Json Mode

Evaluates Groq's Structured Outputs & JSON Mode across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Fast Inference eval coverage.

Mapped capabilities

9 scenarios

  • response_format json_object basics
  • json_schema strict adherence
  • parse safety on malformed JSON

Frequently asked questions

What do the Corsac evals for Groq test?+

Each eval pack tests Groq's public product surface — including Auth Rate Limits And Tiers, Batch Api, and Chat Completions Openai Compatible — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Groq evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Groq cases — from Safety Models And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Groq library.

How many test cases does the Groq library include?+

The Groq eval library includes 73 graded test cases across 8 eval packs, the largest being Safety Models And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Groq or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Groq packs — Auth Rate Limits And Tiers and Batch Api and the rest — against Groq or your own agent with your own data.