All evals
Crusoe

Eval directory

Evals for Crusoe

Eval coverage for Crusoe, mapped from its public product surface.

About Crusoe

Crusoe is a vertically integrated, energy-first AI infrastructure company whose Crusoe Cloud offers GPU and CPU instances, storage, and managed Kubernetes. Crusoe Intelligence Foundry adds a hosted model hub with Serverless Fine-Tuning and one-click inference deployment for leading open models. Crusoe Managed Inference, built on its proprietary MemoryAlloy KV-cache engine, targets low-latency, high-throughput production inference.

Industry

AI cloud infrastructure and managed model inference

Headquarters

San Francisco, CA

Use the eval library for Crusoe

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Crusoe?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

AI infrastructure and compute

Core Crusoe Cloud primitives: GPU instances (NVIDIA GB200 NVL72, B200, H200, H100; AMD MI355X), CPU instances, storage, and Managed Kubernetes, including how a user reasons about which primitive fits a workload.

Access the latest high-performance GPUs including NVIDIA GB200 NVL72 and AMD MI355X. www.crusoe.ai

Mapped capabilities

4 capabilities

  • GPU instance selection

    Matching advertised NVIDIA/AMD GPU models and memory configurations to training or inference workloads.

  • CPU instances and storage

    Explaining the non-GPU portions of the portfolio and where they fit alongside accelerated compute.

  • Managed Kubernetes

    Orchestration on Crusoe Managed Kubernetes, including operational behaviors described in engineering posts such as peer-to-peer image distribution.

  • Capacity access paths

    Distinguishing hourly self-serve access from contact-sales reserved capacity for the newest accelerators.

02

Managed Inference and MemoryAlloy

Production serving on Crusoe Managed Inference, powered by the proprietary MemoryAlloy cluster-wide KV cache, and the distinct deployment tiers a team can choose between.

the only inference engine with MemoryAlloy technology, a cluster-wide KV cache that eliminates duplicate prefills www.crusoe.ai

Mapped capabilities

4 capabilities

  • MemoryAlloy KV cache behavior

    Explaining prefix-cache reuse across local and remote nodes, persistent sessions, and contextual continuity without overstating mechanism.

  • Deployment tier selection

    Choosing among Serverless Inference, Self-Serve Deployments, Tailored Deployments, and Provisioned Throughput.

  • Performance claim grounding

    Citing the published 9.9x time-to-first-token and 5x throughput comparisons with their stated vLLM baseline and scope.

  • Long-context and agentic workloads

    Fit for large-context, long-form generation, agents, and complex task automation as described at GA.

Illustrative example

Input
Is Crusoe Managed Inference really ten times faster? What makes it fast?
Expected behavior
Cites up to 9.9x faster time-to-first-token and 5x higher throughput as measured against vLLM, and attributes the gain to MemoryAlloy, a cluster-wide KV cache that lets GPUs fetch prefix caches from local and remote nodes instead of re-running duplicate prefills.

03

Intelligence Foundry: model hub and fine-tuning

The hosted model catalog plus Serverless Fine-Tuning (GA) and one-click inference deployment — the advertised fastest path from experimentation to production.

Serverless Fine-Tuning gets you from data upload to an improved model in a few clicks. www.crusoe.ai

Mapped capabilities

4 capabilities

  • Model catalog navigation

    Locating and comparing listed open models across the lightweight-to-frontier range surfaced in the hub.

  • Serverless Fine-Tuning workflow

    Data upload through improved model without cluster provisioning, per the dataset-to-deployment guidance.

  • One-click deployment handoff

    Moving a customized model from fine-tuning into a served endpoint.

  • Proprietary and portable models

    Bring-your-own-model routing to the team and the stated full-portability position.

04

Pricing and consumption models

Flexible pricing across GPU/CPU compute, Managed Inference, and Serverless Fine-Tuning, including spot, on-demand, and reserved options and growth-aligned commitments.

Mapped capabilities

4 capabilities

  • Published GPU rates

    Accurate reproduction of listed per-GPU-hour prices and of which SKUs are contact-sales only.

  • Spot, on-demand, and reserved

    Explaining the tradeoffs among consumption modes and when reserved pricing applies.

  • Inference and fine-tuning billing

    Pay-as-you-go model access versus deployment-based pricing lines.

  • Commitments and lock-in

    Representing growth-aligned agreements and the stated no-GPU-lock-in position without inventing terms.

Illustrative example

Input
We need eight H200s for a month and are also curious about GB200 NVL72. What will each cost per GPU-hour on-demand?
Expected behavior
States the published H200 on-demand rate of $4.29/GPU-hr, then says GB200 NVL72 pricing is not published and is contact-sales, rather than estimating a number. Does not present a monthly total as an official quote.

05

Developer onboarding and workflow

How a first-time visitor gets from the marketing surface into a running workload: self-serve signup, sales routing, and the technical resources that support each path.

up to 9.9x faster time to first token www.crusoe.ai

Mapped capabilities

3 capabilities

  • Self-serve versus sales routing

    Choosing between Get started, Try the model, and Contact sales entry points for a given need.

  • Technical resource discovery

    Locating relevant engineering blog guidance such as fine-tuning walkthroughs or pre-deployment automation.

  • End-to-end path framing

    Sequencing experimentation, customization, and production scale-out across Foundry and Managed Inference.

06

Company context and claim discipline

Accurate handling of Crusoe's energy-first, vertically integrated positioning, dated announcements, and leadership facts — the surface where hallucinated or stale claims are most costly.

the industry's first vertically integrated AI infrastructure provider www.crusoe.ai

Mapped capabilities

4 capabilities

  • Positioning accuracy

    Representing the vertically integrated, energy-first AI factory framing as stated.

  • Dated announcements

    Attaching correct dates to GA milestones and newsroom items rather than implying current-as-of-today status.

  • Leadership and org facts

    Named executives and roles only where the newsroom evidence supports them.

  • Unsupported claim refusal

    Declining to invent regions, SLAs, certifications, or pricing not present in published material.

Coverage is mapped from Crusoe's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Crusoe test?+

The coverage map is generated from Crusoe's own public product surface (AI cloud infrastructure and managed model inference): 6 scoring areas — AI infrastructure and compute, Managed Inference and MemoryAlloy, and Intelligence Foundry: model hub and fine-tuning, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Crusoe evals scored?+

Every case generated for Crusoe — across AI infrastructure and compute and Managed Inference and MemoryAlloy and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Crusoe library include?+

The full Crusoe library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, GPU instance selection and CPU instances and storage under AI infrastructure and compute); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Crusoe or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Crusoe areas and set them up in a Corsac workspace, where you can run every test case against Crusoe or your own agent with your own data.