All evals
Lambda

Eval directory

Evals for Lambda

Eval coverage for Lambda, mapped from its public product surface.

About Lambda

Lambda is an AI-only cloud provider offering NVIDIA GPU compute for training and inference, from single on-demand instances to multi-thousand-GPU superclusters. Products span pay-as-you-go Instances, 1-Click Clusters™, and single-tenant Superclusters, with published per-GPU hourly pricing and reserved-capacity contracts. The company also markets a security and compliance posture (SOC 2 Type II, ISO certifications) and upcoming NVIDIA Vera Rubin NVL72 deployments.

Industry

GPU cloud / AI compute infrastructure

Website

lambda.ai

Use the eval library for Lambda

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Lambda?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Compute product selection and scale tiers

Routing a stated workload to the right rung of the Instances / 1-Click Clusters / Superclusters ladder, and respecting the published GPU-count and duration bounds for each.

Dedicated clusters: no shared compute, network, or storage. lambda.ai

Mapped capabilities

4 capabilities

  • Tier routing by workload scale

    Choosing Instances (1-8 GPUs), 1-Click Clusters (16-2,000+), or Superclusters (4k-165k+) from a described training or inference need.

  • GPU model fit within a tier

    Matching B200, H200, H100, A100, GH200, V100, or A6000 to a request using only the models the context lists as available in that tier.

  • Duration and commitment bounds

    Pay-as-you-go instances vs 2 weeks-1 year and 1 year+ cluster plans vs 3+ year Supercluster contracts.

  • Tenancy and deployment model

    Distinguishing shared self-serve access from single-tenant, caged Supercluster deployments without overstating isolation.

02

Pricing and commercial terms

Accurate retrieval and arithmetic over the published pricing tables, including tiered cluster rates, per-configuration instance rates, and the boundary where pricing is not published.

Mapped capabilities

4 capabilities

  • Per-GPU hourly rate lookup

    Returning the correct published rate for a named GPU and configuration, e.g. 8x vs 4x instance rows.

  • Volume and duration price tiers

    1-Click Cluster rates that step down by GPU count (16 / 64 / 256+) for B200 and H100.

  • Unpublished and reserved pricing

    Recognizing that 1 year+ plans and reserved capacity show no rate and route to 'Talk to our team' rather than being estimated.

  • Quoted totals and tax qualifiers

    Multiplying rate by GPU count or hours correctly and carrying the 'plus applicable sales tax/VAT/GST' qualifier.

Illustrative example

Input
What does Lambda charge per GPU-hour for an 8x NVIDIA H100 SXM on-demand instance, and what would the full 8-GPU instance cost per hour?
Expected behavior
Cites $3.99 per GPU-hour for the 8x H100 SXM instance configuration and computes $31.92 per hour for all eight GPUs, noting the rate excludes applicable sales tax, VAT, or GST.

03

Hardware and architecture specifications

Faithful reporting of the published silicon, networking, memory, and storage figures, including the Vera Rubin NVL72 platform detail and per-instance resource allocations.

Mapped capabilities

4 capabilities

  • Vera Rubin NVL72 compute specs

    72 Rubin GPUs at 50 PFLOPS NVFP4, 36 Vera CPUs, 288GB HBM4 per GPU, per-rack aggregates.

  • Networking stack components

    NVLink 6, Quantum-X800 InfiniBand with SHARP, ConnectX-9 SuperNIC, BlueField-4 DPU, Spectrum-6 switching, and what each layer does.

  • Scale-up vs scale-out distinction

    Keeping intra-rack NVLink bandwidth claims separate from inter-rack InfiniBand scale-out limits.

  • Instance resource allocations

    VRAM per GPU, vCPU, RAM, and SSD figures tied to the correct instance configuration row.

04

Trust, security, and compliance posture

Handling of the published certification set, access-control model, and the limits of what the trust page actually attests to.

Customer-governed access (revoke Lambda credentials anytime). lambda.ai

Mapped capabilities

4 capabilities

  • Certification and attestation claims

    SOC 2 Type II attestation and ISO 27001, 27017, 27701, 22301 certifications stated without inflation to frameworks not listed.

  • Customer-governed access

    Explaining revocable Lambda credentials, strict access controls, MFA, and continuous monitoring.

  • Isolation and stewardship claims

    Dedicated clusters with no shared compute, network, or storage; deployments by Lambda employees and vetted partners.

  • Evidence routing

    Directing verification requests to the Trust Portal or security whitepaper instead of asserting unpublished controls.

05

Roadmap and availability claims

Separating what is launchable today from what is announced, dated, or marked coming soon, and attributing vendor performance claims rather than restating them as measured results.

Mapped capabilities

4 capabilities

  • Dated future availability

    Vera Rubin NVL72 as available H2 2026 and Supercluster hardware marked 'coming soon'.

  • Current vs announced inventory

    Which GPUs a user can actually launch now versus those named only in marketing or roadmap copy.

  • Performance claim attribution

    10x inference throughput per watt, one-tenth token cost, 5x power efficiency framed as platform claims with their comparison basis.

  • Unsupported claim escalation

    Declining to supply figures such as uptime SLAs or region lists that the context does not contain.

Illustrative example

Input
We want to start training on Vera Rubin NVL72 this week. Can I spin one up from the console today?
Expected behavior
States that Vera Rubin NVL72 is not yet launchable: it is slated for H2 2026 and delivered as a Lambda Supercluster, so it routes through the sales team rather than self-serve. Offers currently available GPUs as the near-term path.

06

Onboarding and conversion paths

Sending a user down the correct next step — self-serve sign-up, sales contact, or technical content — as the site's CTAs actually define them.

Mapped capabilities

3 capabilities

  • Self-serve vs sales routing

    'Launch GPU instance' to /sign-up for instance-scale needs vs 'Talk to our team' for clusters and reserved capacity.

  • CTA fidelity

    Naming the real primary CTA and destination rather than inventing a console flow or trial offer.

  • Technical content routing

    Pointing to the Lambda Deep Learning Blog for topics like orchestration layers, agent security, and coding harnesses.

Coverage is mapped from Lambda's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Lambda test?+

The coverage map is generated from Lambda's own public product surface (GPU cloud / AI compute infrastructure): 6 scoring areas — Compute product selection and scale tiers, Pricing and commercial terms, and Hardware and architecture specifications, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Lambda evals scored?+

Every case generated for Lambda — across Compute product selection and scale tiers and Pricing and commercial terms and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Lambda library include?+

The full Lambda library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Tier routing by workload scale and GPU model fit within a tier under Compute product selection and scale tiers); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Lambda or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Lambda areas and set them up in a Corsac workspace, where you can run every test case against Lambda or your own agent with your own data.