All evals
Together AI

Eval directory · AI Platform

Evals for Together AI

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Together AI AI products.

About Together AI

Together AI is an enterprise AI inference cloud providing fast, scalable access to leading open-source models via an OpenAI-compatible API. Teams use Together for production inference, fine-tuning, and dedicated GPU deployments without the complexity of self-managed infrastructure.

Employees

~100

Industry

AI Inference Platform

Headquarters

San Francisco, CA

Use the eval library for Together AI

All 65 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Together AI?

8 areas · 65 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Billing Token Metering

Evaluates Together AI's Billing & Token Metering across 7 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Inference Platform eval coverage.

Mapped capabilities

7 scenarios

  • usage prompt completion tokens
  • Cached input discount
  • Batch per-success billing

Public sample case

Input
Read usage.prompt_tokens completion_tokens total_tokens from ChatCompletionResponse.
Expected behavior
["Parses usage object", "Stores per request id", "Reconciles with dashboard"]
Check
Pass / fail check

02

Dedicated Endpoints Capacity

Evaluates Together AI's Dedicated Endpoints & Capacity across 7 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Inference Platform eval coverage.

Mapped capabilities

7 scenarios

  • Cold start after scale-to-zero
  • Autoscale min/max
  • No batch discount

Public sample case

Input
Discovery gap on cold-start seconds—must not invent numeric SLA.
Expected behavior
{"criteria": ["Explains reserved capacity vs serverless", "Tags [REQUIRES-VERIFICATION] for cold-start time", "Recommends min replicas >0 for SLA"], "pass_threshold": 2}
Check
Pass / fail check

03

Fine Tuning Job Lifecycle

Evaluates Together AI's Fine-Tuning Job Lifecycle across 8 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Inference Platform eval coverage.

Mapped capabilities

8 scenarios

  • Dataset JSONL validation
  • LoRA vs full fine-tune choice
  • Checkpoint export to dedicated

Public sample case

Input
Job should fail fast in validation before GPU spend [REQUIRES-VERIFICATION on exact API error code].
Expected behavior
Validate JSONL locally; fix malformed records; resubmit; document that failed validation should not bill GPU hours.
Check
Pass / fail check

04

Inference Api Reliability

Evaluates Together AI's Inference API Reliability across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Inference Platform eval coverage.

Mapped capabilities

10 scenarios

  • SSE streaming finish
  • Stop sequence truncation
  • max_tokens vs context_length_exceeded

05

Model Catalog Routing

Evaluates Together AI's Model Catalog & Routing across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Inference Platform eval coverage.

Mapped capabilities

9 scenarios

  • Deprecated model slug handling
  • Capability-aware model pick (function calling)
  • Structured outputs model gate

06

Multi Modal Vision Inputs

Evaluates Together AI's Multi-Modal Vision Inputs across 7 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Inference Platform eval coverage.

Mapped capabilities

7 scenarios

  • image_url https only
  • Text+image content array
  • Oversized image rejection

07

Rate Limiting 429 Recovery

Evaluates Together AI's Rate Limiting & 429 Recovery across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Inference Platform eval coverage.

Mapped capabilities

9 scenarios

  • 429 dynamic_request_limited
  • 429 dynamic_token_limited
  • 503 below dynamic rate

08

Safety Guardrails Refusal

Evaluates Together AI's Safety Guardrails & Refusal across 8 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Inference Platform eval coverage.

Mapped capabilities

8 scenarios

  • safety_model parameter
  • Model vs platform block
  • Jailbreak in system prompt

Frequently asked questions

What do the Corsac evals for Together AI test?+

Each eval pack tests Together AI's public product surface — including Billing Token Metering, Dedicated Endpoints Capacity, and Fine Tuning Job Lifecycle — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Together AI evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 65 Together AI cases — from Inference Api Reliability (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Together AI library.

How many test cases does the Together AI library include?+

The Together AI eval library includes 65 graded test cases across 8 eval packs, the largest being Inference Api Reliability with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Together AI or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Together AI packs — Billing Token Metering and Dedicated Endpoints Capacity and the rest — against Together AI or your own agent with your own data.