All evals
CoreWeave

Eval directory

Evals for CoreWeave

Eval coverage for CoreWeave, mapped from its public product surface.

About CoreWeave

CoreWeave is a cloud platform purpose-built for AI workloads, offering NVIDIA GPU compute, purpose-built storage, high-performance InfiniBand networking, and managed software services in a Kubernetes-native environment. It includes SUNK, a unified Slurm-on-Kubernetes training system, plus Mission Control for cluster health monitoring and observability. Pricing is published per GPU-hour with on-demand and spot options, and the company positions itself on lower total cost of ownership versus general-purpose clouds.

Industry

AI cloud infrastructure (GPU compute, storage, and networking)

Headquarters

Livingston, NJ

Use the eval library for CoreWeave

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for CoreWeave?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

GPU and Bare Metal Compute

Selecting and describing the right compute shape for an AI workload across the published NVIDIA GPU lineup, CPU compute, and bare metal, including the tradeoffs of virtualization-free access and DPU offload.

Get the GPU compute you need for your AI workloads though a Kubernetes-native environment www.coreweave.com

Mapped capabilities

4 capabilities

  • GPU model and instance spec lookup

    VRAM, vCPU, system RAM, local storage, and GPU count per published instance type

  • Bare metal vs. Kubernetes-native tradeoffs

    When full-access bare metal, x86/Arm CPUs, or BlueField DPU offload apply

  • Workload-to-hardware fit

    Matching training vs. inference workloads to Blackwell, GB200/GB300, or RTX PRO options

  • Capacity and availability qualifiers

    Handling 'contact sales' tiers and region-scoped listings without inventing availability

02

Pricing, Billing, and TCO

Reasoning over published per-GPU-hour on-demand, spot, and single-GPU inference pricing, and over the vendor's comparative total-cost-of-ownership positioning against general-purpose clouds.

Up to 47 % Lower total cost over 3 years www.coreweave.com

Mapped capabilities

4 capabilities

  • Published rate retrieval

    On-demand, spot, and inference single-GPU hourly prices for listed instances

  • Cost estimation and arithmetic

    Multi-hour or multi-node run cost from published rates

  • TCO claim attribution

    Citing the 3-year analysis figures as vendor claims with stated scope, not universal truths

  • Fee structure and hidden-cost handling

    Egress, API request, and orchestration charge claims; usage-based storage tiering

Illustrative example

Input
What does one NVIDIA HGX B200 8-GPU node cost per hour on-demand versus spot on CoreWeave, and what would 24 hours on-demand cost?
Expected behavior
States $68.80/hour on-demand and $34.11/hour spot for the 8-GPU HGX B200 node, and computes $1,651.20 for 24 on-demand hours. Notes prices are the published North America list rates.

03

Storage and Data Path

Explaining purpose-built AI storage: throughput to GPUs, caching, checkpointing, snapshot and retention policy, security posture, and separation of storage from compute.

Automated snapshots occur every 6-hour interval with a 3-day retention www.coreweave.com

Mapped capabilities

4 capabilities

  • LOTA caching behavior

    On-node caching and the published per-GPU throughput figure

  • Snapshot, retention, and checkpointing

    Automated snapshot interval, retention window, and resuming after interruption

  • Automated tiering and billing model

    Usage-based movement of inactive data without manual tiering or multiple APIs

  • Storage security posture

    Encryption at rest and in transit, IAM, authentication, and role-based policy

Illustrative example

Input
How often does CoreWeave Storage take automated snapshots, and how long are they retained?
Expected behavior
States that automated snapshots occur every 6 hours with a 3-day retention period. Does not invent additional tiers, configurable schedules, or cross-region replication guarantees not present in the published description.

04

Networking and Cluster Scale-Out

Describing the InfiniBand fabric, cluster scale limits, and connectivity options for private, hybrid, and multi-cloud topologies.

Mapped capabilities

4 capabilities

  • InfiniBand fabric characteristics

    SHARP-enabled Quantum InfiniBand, non-blocking channels, latency and per-node throughput claims

  • Supercluster and megacluster scale

    Published GPU-count ceilings for interconnected clusters

  • VPC and traffic isolation

    Private customer traffic and offloading networking from GPU nodes

  • Direct Connect and hybrid links

    Port speeds, dedicated ports vs. carrier connections, on-prem and hyperscaler links

05

SUNK Training Workflow

The unified Slurm-on-Kubernetes surface: onboarding and provisioning, cluster deployment, topology-aware scheduling, and portability of one training workflow across environments.

Mapped capabilities

4 capabilities

  • User provisioning and onboarding

    SUP-driven secure onboarding and identity/config drift reduction

  • Cluster deployment via self-service

    Bringing production-ready SUNK clusters online and managing them

  • Scheduling for distributed jobs

    Topology-aware scheduling and shared training/inference pod scheduling

  • Portability with SUNK Anywhere

    Preserving one way of running workloads beyond CoreWeave infrastructure

06

Cluster Health, Observability, and Recovery

Mission Control and platform observability: detecting silent hardware faults and stragglers, correlating job and infrastructure signals, and recovering long-running training runs from disruption.

Mapped capabilities

4 capabilities

  • Fault and straggler detection

    Silent hardware issues and GPU stragglers surfaced before they compound

  • Cross-layer signal correlation

    Slurm job metrics against GPU, network, and storage telemetry

  • Interruption recovery workflow

    Resuming multi-day runs from checkpoints after hardware interruption

  • Utilization and goodput reporting

    Framing MFU and goodput figures as vendor-published comparisons

Coverage is mapped from CoreWeave's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for CoreWeave test?+

The coverage map is generated from CoreWeave's own public product surface (AI cloud infrastructure (GPU compute, storage, and networking)): 6 scoring areas — GPU and Bare Metal Compute, Pricing, Billing, and TCO, and Storage and Data Path, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the CoreWeave evals scored?+

Every case generated for CoreWeave — across GPU and Bare Metal Compute and Pricing, Billing, and TCO and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the CoreWeave library include?+

The full CoreWeave library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, GPU model and instance spec lookup and Bare metal vs. Kubernetes-native tradeoffs under GPU and Bare Metal Compute); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against CoreWeave or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped CoreWeave areas and set them up in a Corsac workspace, where you can run every test case against CoreWeave or your own agent with your own data.