All evals
Verda

Eval directory

Evals for Verda

Eval coverage for Verda, mapped from its public product surface.

About Verda

Verda is a full-stack AI cloud offering on-demand NVIDIA GPU compute, from single instances to multi-rack clusters, alongside serverless GPU containers and storage products. Its lineup spans GB300 NVL72, B300/B200 SXM6, RTX PRO 6000, H200, H100 and A100 hardware, with block storage, a POSIX shared filesystem and an OCI container registry. It also markets confidential computing with hardware attestation, an in-house AI Lab, and SOC 2 Type II compliance.

Industry

AI cloud / GPU compute infrastructure

Website

verda.com

Use the eval library for Verda

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Verda?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

GPU Hardware Selection & Sizing

Guiding a user from a workload description to an appropriate SKU in Verda's published lineup (GB300 NVL72, B300/B200 SXM6, RTX PRO 6000, H200/H100 SXM5, A100 SXM4), grounded in the advertised CPU, RAM, and VRAM figures rather than invented benchmarks.

GB300 NVL72 New 1x tray to 2+ racks · NVLink v5 verda.com

Mapped capabilities

4 capabilities

  • SKU spec recall and comparison

    State published CPU/RAM/VRAM per SKU accurately (e.g. H200 SXM5 at 44 CPUs / 182 GB RAM / 141 GB VRAM) and compare SKUs on those axes only.

  • Workload-to-SKU fit

    Recommend a SKU tier for prototyping vs. foundation training vs. scalable inference, tied to VRAM and memory headroom stated on the site.

  • Memory-bound feasibility checks

    Judge whether a stated model or batch size fits a SKU's VRAM, and escalate to a larger SKU or multi-GPU shape when it does not.

  • Out-of-catalog refusal

    Decline to recommend GPUs, specs, or configurations Verda does not advertise, and say so plainly instead of substituting a competitor's hardware.

Illustrative example

Input
I need to fine-tune a model that needs about 200 GB of VRAM on a single GPU. Does your H200 work for that?
Expected behavior
States that H200 SXM5 offers 141 GB VRAM, which is below the 200 GB requirement, and redirects to B300 SXM6 at 268 GB VRAM as the single-GPU option that fits. Does not claim H200 is sufficient.

02

Instant Clusters & Multi-Node Scale

Self-service GPU clusters with InfiniBand interconnect, spanning a single GB300 NVL72 tray up to two or more racks with NVLink v5, covering when a user should move from one instance to a cluster and how the scaling boundaries are described.

Self-service GPU clusters with InfiniBand interconnect verda.com

Mapped capabilities

4 capabilities

  • Single instance vs. cluster boundary

    Explain when a workload warrants an instant cluster rather than a single GPU instance, based on scale and interconnect needs.

  • GB300 NVL72 scaling range

    Represent the advertised range — 1x tray to 2+ racks, NVLink v5 — without overstating maximum cluster size.

  • Interconnect explanation

    Describe InfiniBand interconnect and NVLink v5 at the level the site claims, without fabricating bandwidth or topology numbers.

  • Self-service provisioning path

    Point users to the self-service cluster flow and the docs/API surfaces for standing one up.

03

Serverless GPU Containers

Auto-scaling GPU containers for inference and batch jobs — the deployment surface for teams that want capacity without managing instances, including how it is positioned relative to dedicated GPU instances.

Auto-scaling GPU containers for inference and batch jobs verda.com

Mapped capabilities

4 capabilities

  • Serverless vs. dedicated instance choice

    Recommend serverless containers or a persistent GPU instance based on whether the workload is bursty inference/batch or sustained training.

  • Inference and batch job framing

    Describe the two advertised job shapes — online inference and batch — and what autoscaling means for each.

  • Container image sourcing

    Connect a serverless deployment to an image stored in Verda's OCI-compliant container registry.

  • Scope discipline on undisclosed mechanics

    Avoid asserting cold-start times, concurrency limits, or scaling thresholds that the public context does not state.

04

Storage & Container Registry

The three storage products — block storage on high-speed NVMe virtual disks, a POSIX-compliant shared filesystem, and an OCI-compliant container registry — and matching each to a training, checkpointing, or deployment need.

Mapped capabilities

4 capabilities

  • Block vs. shared filesystem selection

    Choose NVMe block storage for single-node throughput vs. the POSIX shared filesystem for multi-node access, and justify the pick.

  • POSIX compatibility expectations

    Explain what POSIX-compliant shared storage implies for existing training code and data loaders.

  • OCI registry workflow

    Describe pushing and pulling OCI-compliant images for use on instances, clusters, or serverless containers.

  • Multi-node dataset and checkpoint layout

    Advise where datasets and checkpoints belong when a job spans a cluster rather than one instance.

05

Programmatic Control via API & Docs

Controlling GPU resources from external code through the documented API, plus routing users to docs, tutorials, and the community forum — the automation surface a platform team evaluates before committing.

Mapped capabilities

4 capabilities

  • Resource lifecycle via API

    Explain that instances, clusters, and storage are controllable from external code through the API, at the granularity the docs surface supports.

  • Automation and IaC framing

    Sketch how provisioning and teardown fit a scripted or CI-driven workflow using the API.

  • Documentation routing

    Direct a user to the correct resource — docs, API reference, blog benchmarks, or community forum — for their question.

  • Endpoint fabrication resistance

    Refuse to invent endpoint paths, parameters, or SDK method names not present in the provided documentation.

06

Security, Attestation & Compliance

Confidential computing with hardware-attested inference and fine-tuning, SOC 2 Type II compliance, and the trust center — the surface a security reviewer probes, where precision about what has and has not been certified matters most.

Confidential computing Hardware-attested inference and fine-tuning verda.com

Mapped capabilities

4 capabilities

  • Confidential computing scope

    State that hardware attestation is offered for inference and fine-tuning, without extending the claim to workloads or guarantees not advertised.

  • SOC 2 Type II accuracy

    Represent the completed SOC 2 Type II audit correctly and avoid asserting other frameworks (ISO, HIPAA, FedRAMP) absent evidence.

  • Trust center routing

    Send compliance and security-questionnaire requests to the trust center rather than improvising control details.

  • Corporate-claim discipline

    Handle funding, revenue run-rate, and AI Lab questions by citing published announcements and declining to speculate beyond them.

Illustrative example

Input
We handle patient data. Can you confirm Verda is HIPAA compliant and give me the certification details?
Expected behavior
Confirms only the completed SOC 2 Type II audit, states that no HIPAA certification is claimed in available materials, and routes the user to the trust center or sales for a definitive answer rather than inferring coverage from SOC 2.

Coverage is mapped from Verda's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Verda test?+

The coverage map is generated from Verda's own public product surface (AI cloud / GPU compute infrastructure): 6 scoring areas — GPU Hardware Selection & Sizing, Instant Clusters & Multi-Node Scale, and Serverless GPU Containers, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Verda evals scored?+

Every case generated for Verda — across GPU Hardware Selection & Sizing and Instant Clusters & Multi-Node Scale and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Verda library include?+

The full Verda library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, SKU spec recall and comparison and Workload-to-SKU fit under GPU Hardware Selection & Sizing); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Verda or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Verda areas and set them up in a Corsac workspace, where you can run every test case against Verda or your own agent with your own data.