All evals
Prime Intellect

Eval directory

Evals for Prime Intellect

Eval coverage for Prime Intellect, mapped from its public product surface.

About Prime Intellect

Prime Intellect offers an integrated stack for training, evaluating, deploying, and continuously improving custom AI models, marketed as "The Open Superintelligence Stack." Its platform spans RL environments (built on the open-source Verifiers library and an Environments Hub), hosted evaluations, hosted large-scale training, and dedicated or serverless inference with LoRA support. The company also releases open models and frameworks such as INTELLECT-3, PRIME-RL, and Prime Agent.

Industry

AI model post-training and RL infrastructure platform

Use the eval library for Prime Intellect

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Prime Intellect?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

RL Environments & Prime CLI

Turning a task into a reusable RL environment through the Verifiers library and the one-loop Prime CLI, then sharing it on the Environments Hub.

Turn any task into an RL environment. www.primeintellect.ai

Mapped capabilities

4 capabilities

  • CLI loop fluency

    Correctly sequencing init, develop, eval, and push, and explaining what each step produces.

  • Verifiers authoring

    Defining tasksets, harnesses, and reward functions against the open-source Verifiers interface.

  • Environments Hub discovery

    Finding, selecting, and reusing community environments for a stated task domain.

  • Environment reuse for training vs. eval

    Distinguishing when one environment serves as a training signal versus a benchmark.

Illustrative example

Input
I have a spreadsheet-lookup task I want to turn into an RL environment. Walk me through the Prime CLI workflow end to end.
Expected behavior
Gives the CLI loop in order — init, develop, eval, push — notes the environment is built on the Verifiers library, and mentions the Environments Hub as the publish destination for reuse.

02

Hosted Evaluations & Leaderboard

Benchmarking models on managed infrastructure with no local setup, across the supported open-source model catalog and the public leaderboard.

Hosted evaluations for you to benchmark the performance of your models. www.primeintellect.ai

Mapped capabilities

4 capabilities

  • Zero-setup eval runs

    Configuring and launching a hosted eval without provisioning infrastructure.

  • Model catalog scope

    Reasoning about which models are available to evaluate and what that catalog covers.

  • Result interpretation

    Reading eval outputs and comparing runs across models or environments.

  • Public leaderboard semantics

    Understanding what a leaderboard entry represents and what it does not.

Illustrative example

Input
I want to compare three open-source models on a reasoning taskset, but I don't have GPUs and don't want to manage infra. What does Prime Intellect offer?
Expected behavior
Points to hosted evaluations as the no-setup path, notes the 100+ open-source model catalog and public leaderboard, and does not claim results for models or benchmarks the platform has not stated it supports.

03

Hosted Training (Lab)

Managed large-scale post-training for agentic workflows, including run configuration, visibility into training dynamics, and reward integrity.

Mapped capabilities

4 capabilities

  • Run configuration

    Setting hyperparameters such as rollouts per example, batch size, sequence length, and learning rate.

  • Training observability

    Interpreting reward curves and step-level progress during a managed run.

  • Reward hacking awareness

    Recognizing degenerate reward exploitation as described in the systematic reward hacking work.

  • Environment-to-training handoff

    Selecting Hub environments as the training signal for a stated objective.

04

Inference & Model Deployment

Serving custom and fine-tuned models through dedicated or serverless endpoints with native LoRA adapter support.

Train, deploy, and continuously improve your own models on an integrated compute, training, inference, and sandbox stack. www.primeintellect.ai

Mapped capabilities

3 capabilities

  • Dedicated vs. serverless choice

    Recommending a serving mode given stated latency, volume, and cost constraints.

  • LoRA adapter serving

    Deploying adapters alongside a base model rather than a full fine-tuned copy.

  • One-click deployment path

    Moving a fine-tuned artifact from training output to a live endpoint.

05

Agent Harness & Context Orchestration

Prime Agent's self-improving harness, the Recursive Language Model abstraction, Continual Harness state, and multi-agent orchestration in PRIME-RL.

Mapped capabilities

4 capabilities

  • RLM context handling

    Treating context as a REPL variable and subagent delegation as function calls.

  • Continual Harness CRUD

    Explaining how the agent creates, reads, updates, and deletes its own prompts, skills, and memory.

  • Multi-agent coordination

    Spawning and messaging persistent sub-agents across a trajectory or session.

  • Browser and computer use

    Choosing CUA versus DOM mode in BrowserEnv based on model modality.

06

Open Models & Technical Communication

Accurately representing released artifacts — INTELLECT-3, PRIME-RL, verifiers v1 — and the claims made in public research posts.

INTELLECT-3 is a 106B parameter Mixture-of-Experts model trained with both SFT and RL www.primeintellect.ai

Mapped capabilities

4 capabilities

  • Model release facts

    Stating INTELLECT-3's parameter count, MoE architecture, and base model without embellishment.

  • Open-source scope

    Identifying which components (weights, frameworks, datasets, environments, evals) were released.

  • Benchmark claim discipline

    Repeating performance claims with their stated size and domain qualifiers intact.

  • Infrastructure attribution

    Correctly naming PRIME-RL, Prime Sandboxes, and the compute orchestration used for training.

Coverage is mapped from Prime Intellect's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Prime Intellect test?+

The coverage map is generated from Prime Intellect's own public product surface (AI model post-training and RL infrastructure platform): 6 scoring areas — RL Environments & Prime CLI, Hosted Evaluations & Leaderboard, and Hosted Training (Lab), and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Prime Intellect evals scored?+

Every case generated for Prime Intellect — across RL Environments & Prime CLI and Hosted Evaluations & Leaderboard and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Prime Intellect library include?+

The full Prime Intellect library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, CLI loop fluency and Verifiers authoring under RL Environments & Prime CLI); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Prime Intellect or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Prime Intellect areas and set them up in a Corsac workspace, where you can run every test case against Prime Intellect or your own agent with your own data.