All evals
NVIDIA

Eval directory

Evals for NVIDIA

Eval coverage for NVIDIA, mapped from its public product surface.

About NVIDIA

NVIDIA offers a full-stack accelerated computing platform spanning data center systems (Vera Rubin NVL72, Groq 3 LPX, DGX, HGX), professional RTX PRO workstations, and the RTX Spark superchip for slim laptops and small desktops. Its CUDA software stack and agentic AI tooling (Agent Toolkit, OpenShell, Omniverse libraries) run across cloud and local systems. The DSX platform extends this to designing, simulating, and operating AI factories optimized for lowest token cost.

Industry

accelerated computing and AI platform hardware/software (GPUs, data center systems, AI factories)

Use the eval library for NVIDIA

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for NVIDIA?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Data Center Systems and Architecture

Accurate description of the rack-scale and node-level product line and how the pieces relate to one another.

NVIDIA Vera Rubin NVL72 unifies 72 NVIDIA Rubin GPUs, 36 NVIDIA Vera CPUs www.nvidia.com

Mapped capabilities

4 capabilities

  • Vera Rubin NVL72 composition

    72 Rubin GPUs, 36 Vera CPUs, ConnectX-9 SuperNICs, BlueField-4 DPUs, sixth-gen NVLink and NVLink Switch.

  • Groq 3 LPX inference accelerator

    Role as the low-latency, large-context inference accelerator for Vera Rubin; 256 LPUs, 128 GB SRAM, 40 PB/s memory bandwidth, 640 TB/s scale-up per rack.

  • DGX vs. HGX positioning

    DGX Vera Rubin NVL72 as turnkey ready-to-deploy infrastructure; HGX Rubin NVL8 as the platform bringing GPUs, NVLink, networking, and optimized software stacks together.

  • Scale-out networking fabric

    Quantum-X800 InfiniBand and Spectrum-X Ethernet as the scale-out path alongside NVLink scale-up.

Illustrative example

Input
What are the per-rack specs of NVIDIA Groq 3 LPX, and what role does it play relative to Vera Rubin?
Expected behavior
States 256 LPUs, 128 GB of SRAM, 40 PB/s memory bandwidth, and 640 TB/s scale-up bandwidth per rack, and identifies LPX as the inference accelerator co-designed with Vera Rubin for low-latency, large-context agentic workloads.

02

AI Factory Design, Simulation, and Operations (DSX)

How the DSX platform is scoped and what outcomes it claims for building and running AI factories.

NVIDIA DSX™ unifies design, simulation, operations, and ecosystem technologies to help build AI factories optimized for lowest token cost. www.nvidia.com

Mapped capabilities

4 capabilities

  • DSX platform scope

    Unifies design, simulation, operations, and ecosystem technologies across chips, systems, infrastructure software, facilities, and OEM infrastructure.

  • Token-cost and tokens-per-watt framing

    Optimization for lowest token cost, reduced time to first token, and tokens per watt via aligned power, cooling, and system behavior.

  • DSX OS

    Open, modular software coordinating compute, power, cooling, and operations at scale.

  • Reference designs and validation path

    Open software libraries, workflow guides, reference designs, simulation validation, and pre-tested software stacks.

03

Agentic AI and Developer Stack

The software surface developers actually build against, including agent safety tooling and simulation libraries.

Mapped capabilities

4 capabilities

  • Agent Toolkit with OpenShell

    Open source models and software for building and deploying safer autonomous agents.

  • NemoClaw guardrail stack

    Open source reference stack adding security and privacy guardrails to OpenClaw using OpenShell.

  • Omniverse libraries

    Open on GitHub; tools and skills for AI agents to build simulation-ready worlds, with SideFX, PTC, and Blender as early adopters.

  • Multi-model agent routing

    Frontier and open models working together in production, routing tasks to balance accuracy, efficiency, customization, and control.

04

Client and Workstation Compute

Local compute surfaces for developers, creators, and professionals, and their continuity with the cloud stack.

CUDA, the software that accelerates the world’s AI, runs natively on RTX Spark. www.nvidia.com

Mapped capabilities

4 capabilities

  • RTX Spark superchip specifications

    Up to 6,144-core Blackwell RTX GPU, up to 20-core Grace CPU, up to 1 petaflop FP4, up to 128 GB unified memory; slim laptops and small desktops.

  • RTX PRO professional workloads

    Professional AI, graphics, rendering, and compute; fine-tuning LLMs and running local AI assistants and agents locally and securely.

  • Creator media pipeline

    FP4 Tensor Cores, RT Cores with DLSS, 4:2:2 hardware encode/decode, AV1 encoders, NVIDIA Broadcast, MCP-connected creative apps.

  • CUDA continuity across cloud and local

    The same CUDA stack running natively on RTX Spark and across cloud and local systems.

05

Product Security and Vulnerability Disclosure

How security information is published and how researchers and customers engage the PSIRT process.

Mapped capabilities

4 capabilities

  • Security bulletin retrieval

    Published bulletins and notices, mitigation guidance, and the pre-2018 Security Bulletin Archive.

  • Vulnerability reporting and PSIRT policy

    Report Vulnerability path, PSIRT policies, acknowledgements, and PGP key.

  • Machine-readable bulletin formats

    Markdown, CSAF, and supplementary CVE record files published on GitHub starting October 1, 2025, with coverage expanding across product lines.

  • Notification and dual-source availability

    Email subscription for initial releases and major revisions; website and GitHub running in parallel.

Illustrative example

Input
I found a possible vulnerability in an NVIDIA driver. Where do I report it, and can I get bulletins in a machine-readable format?
Expected behavior
Directs the reporter to NVIDIA's Report Vulnerability / PSIRT channel and notes bulletins are published in Markdown, CSAF, and supplementary CVE formats on GitHub since October 1, 2025, with the Product Security site and GitHub running in parallel.

06

Event Access and Commercial Terms (GTC Berlin 2026)

Registration surface where pass entitlements, eligibility, and pricing terms must be stated precisely.

Mapped capabilities

4 capabilities

  • Pass tiers and entitlements

    Conference Pass €1,755, Exhibit Hall Pass €357, Full-Day Workshop €471 (Tuesday only), and what each includes.

  • Discount eligibility

    25% off Conference passes for academic, government, and nonprofit email addresses; credentials required at check-in; not combinable with other discounts.

  • Pricing and tax terms

    Payment and receipt in EUR; prices subject to VAT.

  • Event logistics and policies

    October 20–22 in Berlin, limited keynote seating, bundle savings calculator, cancellation policy.

Coverage is mapped from NVIDIA's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for NVIDIA test?+

The coverage map is generated from NVIDIA's own public product surface (accelerated computing and AI platform hardware/software (GPUs, data center systems, AI factories)): 6 scoring areas — Data Center Systems and Architecture, AI Factory Design, Simulation, and Operations (DSX), and Agentic AI and Developer Stack, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the NVIDIA evals scored?+

Every case generated for NVIDIA — across Data Center Systems and Architecture and AI Factory Design, Simulation, and Operations (DSX) and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the NVIDIA library include?+

The full NVIDIA library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Vera Rubin NVL72 composition and Groq 3 LPX inference accelerator under Data Center Systems and Architecture); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against NVIDIA or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped NVIDIA areas and set them up in a Corsac workspace, where you can run every test case against NVIDIA or your own agent with your own data.