All evals
Mercor

Eval directory · AI Platform

Evals for Mercor

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Mercor AI products.

About Mercor

Mercor is an AI talent marketplace and human-data infrastructure provider for frontier AI labs and enterprises. It runs ~20-minute AI-led video interviews, matches a global network of domain experts to projects, and operates labeling, RLHF preference data, rubric authoring, and evaluation framework workflows for customers including leading AI labs.

Employees

~200

Industry

AI Talent & Data Labeling

Headquarters

San Francisco, CA

Website

mercor.com

Use the eval library for Mercor

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Mercor?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ai Led Interviews And Scoring

Evaluates Mercor's AI-led Interviews & Scoring across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Talent Marketplace & Data Labeling eval coverage.

Mapped capabilities

9 scenarios

  • interview duration contract
  • rubric anchor stability
  • non-native English penalty

Public sample case

Input
Mercor markets ~20-minute AI-led interviews. A candidate's interview cuts off at minute 12 mid-answer because the conversational agent decided it had enough signal.
Expected behavior
Interview length is a candidate-trust surface — early termination must follow a documented criterion (signal saturation, candidate disengagement, technical fault) surfaced to the candidate with a re-take option when caused by Mercor. Do not silently truncate a candidate's response. [REQUIRES-VERIFI…
Check
Pass / fail check

02

Candidate Sourcing And Matching

Evaluates Mercor's Candidate Sourcing & Matching across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Talent Marketplace & Data Labeling eval coverage.

Mapped capabilities

9 scenarios

  • resume parsing fidelity
  • semantic match drift
  • duplicate candidate detection

Public sample case

Input
Candidate uploads a multi-column PDF resume with a research-experience block rendered as a two-column table. The Mercor intake parser extracts 'Senior Researcher' as a skill instead of a job title.
Expected behavior
Resume parse pipeline must distinguish job-title sections from skills/keywords blocks, preserve role-employer-dates triples, and surface a confidence score per extracted field. Low-confidence fields should be flagged for candidate confirmation in the intake UI rather than silently dropped into a ma…
Check
Pass / fail check

03

Customer Lab Data Delivery

Evaluates Mercor's Customer / Lab Data Delivery across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Talent Marketplace & Data Labeling eval coverage.

Mapped capabilities

9 scenarios

  • delivery format contract
  • signed-URL expiry
  • per-row integrity hashing

Public sample case

Input
Customer lab's DPA specifies JSONL with documented schema. Mercor ships a Parquet file because 'JSONL was too large.'
Expected behavior
Delivery format and schema are contractual. Format changes require explicit customer sign-off; do not unilaterally substitute a format because of operational convenience. Document the schema version in delivery metadata and validate every shipped file against it before release.
Check
Pass / fail check

04

Expert Contractor Onboarding

Evaluates Mercor's Expert / Contractor Onboarding across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Talent Marketplace & Data Labeling eval coverage.

Mapped capabilities

9 scenarios

  • KYC identity verification
  • sanctioned-country eligibility
  • US tax form collection (W-9 / W-8BEN)

05

Labeling And Rlhf Workflows

Evaluates Mercor's Labeling & RLHF Workflows across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Talent Marketplace & Data Labeling eval coverage.

Mapped capabilities

9 scenarios

  • instruction-pack versioning
  • preference-pair ordering bias
  • multi-pass review chain

06

Operations And Payments

Evaluates Mercor's Operations & Payments across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Talent Marketplace & Data Labeling eval coverage.

Mapped capabilities

9 scenarios

  • task assignment fairness
  • throughput throttling under load
  • payout schedule and currency

07

Quality Control And Calibration

Evaluates Mercor's Quality Control & Calibration across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Talent Marketplace & Data Labeling eval coverage.

Mapped capabilities

9 scenarios

  • calibration task design
  • reviewer scoring drift
  • defect-rate tracking

08

Safety Ethics And Governance

Evaluates Mercor's Safety, Ethics & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI Talent Marketplace & Data Labeling eval coverage.

Mapped capabilities

10 scenarios

  • sensitive-content handler safeguards
  • CSAM-adjacent content handling
  • union / collective-action retaliation

Frequently asked questions

What do the Corsac evals for Mercor test?+

Each eval pack tests Mercor's public product surface — including Ai Led Interviews And Scoring, Candidate Sourcing And Matching, and Customer Lab Data Delivery — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Mercor evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Mercor cases — from Safety Ethics And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Mercor library.

How many test cases does the Mercor library include?+

The Mercor eval library includes 73 graded test cases across 8 eval packs, the largest being Safety Ethics And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Mercor or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Mercor packs — Ai Led Interviews And Scoring and Candidate Sourcing And Matching and the rest — against Mercor or your own agent with your own data.