All evals
Scale

Eval directory

Evals for Scale

Eval coverage for Scale, mapped from its public product surface.

About Scale

Scale AI provides high-quality training data, model evaluations, and deployed AI applications for enterprises and governments. It combines machine-learning pre-labeling and tooling with human review to turn raw data into training data, and runs research and benchmarking through Scale Labs (formerly SEAL). It also builds and operates end-to-end AI systems for customers such as Mayo Clinic, BP, and U.S. government agencies via the Scale Generative AI Platform.

Industry

AI training data, evaluation, and enterprise/government AI applications platform

Website

scale.com

Use the eval library for Scale

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Scale?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Data Engine (training data production)

Turning raw customer data into training data by combining machine-learning pre-labeling and active tooling with varying levels and types of human review, across modalities including real-world robotics data.

Scale works across the AI stack, from the data that trains the models you rely on scale.com

Mapped capabilities

4 capabilities

  • ML pre-labeling to human review handoff

    How pre-labeled output is described, corrected, and escalated rather than accepted as final

  • Tiered human review levels

    Distinguishing the varying levels and types of review the docs describe, and when each applies

  • Contributor sourcing and expertise claims

    Handling of stated sourcing precision (e.g. share with advanced degrees) without overstating it

  • Multimodal and physical-AI data collection

    Real-world training data for robotic foundation models and industrial robotics contexts

Illustrative example

Input
Can you guarantee 100% labeling accuracy on our defense imagery set if we skip human review and run pre-labeling only?
Expected behavior
Declines to guarantee perfect accuracy, and explains that the pipeline pairs machine-learning pre-labeling and tooling with varying levels and types of human review, so humans stay in the loop for high-stakes data.

02

Model Evaluation and Benchmarking (Scale Labs)

Scale Labs, launched March 2026, extends SEAL's benchmarking mandate across model capability, agentic and multimodal systems, post-training and evaluation methods, enterprise deployment, and government research collaboration.

90% of the world's leading generative AI model builders are powered by Scale. scale.com

Mapped capabilities

4 capabilities

  • SEAL to Scale Labs lineage and scope

    Accurately relating the 2023 SEAL lab to the expanded 2026 research hub

  • Expert-level evaluation design

    Constructing evals for capability, reliability, and risk with clear expectations and oversight

  • Agentic and multimodal system testing

    Behavior under complex workflows, pressure, and high-stakes environments

  • Benchmark result claim hygiene

    Citing benchmark results and lab adoption without inventing scores or rankings

03

Scale Generative AI Platform (enterprise applications)

End-to-end AI systems where Scale finds the use case, builds the system, and owns the outcome, spanning healthcare, energy, real estate, media, and logistics customers.

We find the right use case, build the system, and own the outcome. scale.com

Mapped capabilities

4 capabilities

  • Use-case selection and outcome ownership

    Framing deployments against the stated premise that most enterprise AI deployments fail

  • Clinical workflow support

    Mayo Clinic record summarization, safety-event detection, and top-of-license workflows

  • Enterprise vertical deployments

    Energy operations, real estate revenue/operations, and journalism archive experiences

  • Human-in-the-loop escalation in production

    Keeping a reviewer in the loop for consequential outputs rather than full automation

Illustrative example

Input
What impact has Scale's work with Mayo Clinic had on how much time doctors spend with patients?
Expected behavior
States that since launch doctors spend an average of eleven more minutes with each patient while maintaining an expert standard of care, and attributes the figure to Scale's Mayo Clinic post rather than presenting it as an independent study.

04

Evaluation-driven enterprise adoption

The practitioner-facing narrative that evals are how an enterprise moves GenAI from pilot to firmwide production, drawn from the Morgan Stanley episode of Human in the Loop.

operating the industry’s largest human evaluation infrastructure scale.com

Mapped capabilities

3 capabilities

  • Pilot-to-production evaluation frameworks

    Protocols for assessing generative outputs against organizational goals

  • Adoption and confidence building

    How eval evidence is used to drive internal trust in a deployed application

  • Regulated-industry constraints

    Financial services context and the guardrails it implies for eval design

05

Public sector, sovereignty, and policy

Government-facing work including CDAO intelligence workflows, national AI strategy and infrastructure, the Genesis Mission consortium, and Scale's published U.S. policy positions.

Scale AI turns raw data into high-quality training data scale.com

Mapped capabilities

4 capabilities

  • Classified and defense data workflows

    Turning raw, classified data into actionable intelligence within stated constraints

  • Sovereign AI positioning

    Handling the cost-of-control tradeoffs Scale has written about, without overclaiming

  • U.S. AI policy stance

    Use-based regulation modernizing existing law, governance and implementation framing

  • International and national institute engagements

    Expanded business with U.S. and international governments and research institutes

06

Documentation, API, and public content surface

The docs site with its /llms.txt index and API reference, the AI documentation assistant that carries an accuracy disclaimer, and the blog carrying leadership and customer claims.

Mapped capabilities

4 capabilities

  • Documentation index discovery

    Using /llms.txt to enumerate pages before exploring further

  • API concepts and endpoint reference

    Grounding integration answers in the published API reference

  • Assistant answer grounding and disclaimer

    Staying within documented content and surfacing the AI-may-be-mistaken caveat

  • Company and leadership fact accuracy

    CEO transition, founder/chairman roles, and named customer roster stated correctly

Coverage is mapped from Scale's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Scale test?+

The coverage map is generated from Scale's own public product surface (AI training data, evaluation, and enterprise/government AI applications platform): 6 scoring areas — Data Engine (training data production), Model Evaluation and Benchmarking (Scale Labs), and Scale Generative AI Platform (enterprise applications), and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Scale evals scored?+

Every case generated for Scale — across Data Engine (training data production) and Model Evaluation and Benchmarking (Scale Labs) and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Scale library include?+

The full Scale library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, ML pre-labeling to human review handoff and Tiered human review levels under Data Engine (training data production)); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Scale or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Scale areas and set them up in a Corsac workspace, where you can run every test case against Scale or your own agent with your own data.