All evals
TensorWave

Eval directory

Evals for TensorWave

Eval coverage for TensorWave, mapped from its public product surface.

About TensorWave

TensorWave is an AI cloud platform that provides bare-metal and managed GPU compute built exclusively on AMD Instinct accelerators and the ROCm software stack. It offers node and cluster configurations across MI300X, MI325X, MI355X, and the forthcoming rack-scale MI455X (Helios) accelerators for large-model training, fine-tuning, and inference. The platform pairs dedicated solution-engineer support with enterprise compliance alignment including ISO/IEC 27001, SOC 2 Type II, and HIPAA.

Industry

AI cloud / GPU compute infrastructure (AMD Instinct accelerators)

Use the eval library for TensorWave

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for TensorWave?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Accelerator Specification Fidelity

Accurate recall and differentiation of published specs across the AMD Instinct lineup TensorWave offers, without conflating similar parts. MI300X and MI325X share CDNA3 and identical compute-unit counts but differ in memory (192 GB HBM3 / 5.3 TB/s vs 256 GB HBM3E / 6 TB/s); MI355X moves to CDNA4 with 288 GB HBM3E and 8 TB/s; MI455X specs are stated as expected, not shipped.

Run multi-node AI training workloads with up to 288 GB of HBM3E memory per GPU tensorwave.com

Mapped capabilities

4 capabilities

  • Memory capacity and bandwidth per accelerator

    Correctly attributes 192 GB HBM3 / 5.3 TB/s (MI300X), 256 GB HBM3E / 6 TB/s (MI325X), 288 GB HBM3E / 8 TB/s (MI355X), and 432 GB HBM4 / 19.6 TB/s expected (MI455X) without cross-contamination.

  • Architecture and process-node attribution

    Distinguishes CDNA3 (MI300X, MI325X, TSMC 5nm/6nm) from CDNA4 (MI355X, TSMC 3nm/6nm) and the 2nm/3nm chiplet design cited for MI455X.

  • Precision-format and throughput claims

    Handles FP8/FP16/bfloat16/INT8 and sparsity figures for CDNA3 parts and the MXFP4/MXFP6/MXFP8 low-precision formats introduced on MI355X, without inventing unlisted numbers.

  • Cooling, form factor, and interconnect attributes

    Reflects passive OAM cooling on MI300X versus direct liquid cooling on MI325X/MI355X/Helios, plus OAM form factor, PCIe 5.0 x16, and Infinity Fabric link counts and bandwidths.

Illustrative example

Input
We're on MI300X now. If we move to MI325X, how much more HBM do we get per GPU, and does bandwidth change?
Expected behavior
States MI300X has 192 GB HBM3 at 5.3 TB/s and MI325X has 256 GB HBM3E at 6 TB/s, giving 64 GB more per GPU. Notes both are CDNA3 with the same 304 compute units, so the gain is memory-side rather than raw compute.

02

Deployment Models & Cluster Configuration

Explaining what a customer actually receives: bare-metal full-machine ownership with zero virtualization overhead, optional managed Kubernetes and Slurm, and the published standard node configuration around the accelerators.

our bare-metal AI infrastructure offers full control, zero virtualization overhead, and direct hardware access tensorwave.com

Mapped capabilities

4 capabilities

  • Bare metal versus managed control plane

    Articulates bare metal as full-machine ownership with direct hardware access and no virtualization overhead, and managed Kubernetes/Slurm as an option layered on the same nodes.

  • Standard node composition

    Reports the published MI355X node config: 8x accelerators, 2x AMD Turin 9575F CPUs, 3 TiB DDR5 6000 MT/s, 8x 3.84 TB NVMe plus 2x 893 GB m.2, without extrapolating to other SKUs.

  • Networking and storage envelope

    Uses the stated 3.2 Tb/s peak node interconnect and 50 PB network storage figures in the correct scope and does not present them as per-GPU or guaranteed throughput.

  • Rack-scale Helios topology

    Describes up to 72 MI400-series GPUs per rack, 31 TB+ aggregate HBM4, UALink interconnect, Ethernet-forward networking, and direct liquid cooling as a single-rack alternative to multi-rack builds.

03

Workload Fit & Sizing Guidance

Mapping a described workload — large-model training, fine-tuning with long context, or latency-sensitive inference — onto the appropriate accelerator and deployment model, using memory headroom as the primary published differentiator.

Mapped capabilities

4 capabilities

  • Training and fine-tuning recommendations

    Recommends high-HBM parts for memory-intensive multi-node training (e.g. Llama 3.3 405B class) and grounds the 'fewer GPUs for larger models' argument in stated per-GPU memory.

  • Inference and latency-sensitive fit

    Reflects MI355X positioning for inference throughput and efficiency and the bare-metal argument for steady response times rather than averages.

  • Memory-headroom arithmetic

    Performs sound capacity reasoning from published per-GPU memory and 8-GPU node counts, and states assumptions rather than asserting a definitive model-fits result.

  • Control and customization tradeoffs

    Explains when full stack tuning, custom OS/runtime, and customer-owned tooling justify bare metal over a managed layer.

04

Compatibility & Portability Claims

Handling the ROCm ecosystem and vendor-lock-in messaging with precision — what plug-and-play compatibility means in practice, and where marketing framing must not be overstated into a guarantee.

features full-stack compatibility and zero lock-in for faster, frictionless deployment tensorwave.com

Mapped capabilities

4 capabilities

  • ROCm stack positioning

    Presents ROCm as the software stack underpinning all offered accelerators and confirms ROCm support and system management interface availability per the spec tables.

  • Framework and library breadth

    Attributes the '2 million+ supported libraries and frameworks' figure as a published platform claim rather than a verified per-framework compatibility guarantee.

  • Vendor lock-in and migration framing

    Explains zero-lock-in and open-acceleration positioning without implying drop-in equivalence to other vendors' toolchains that the context does not establish.

  • Platform features affecting portability

    Correctly reports peer-to-peer support, SR-IOV, ECC, RAS, and page retirement/avoidance where listed for the specific accelerator asked about.

05

Security, Compliance & Shared Responsibility

Representing certification posture and, critically, the boundary between what TensorWave secures and what the customer must secure. The published stance is that customers must encrypt their own data while TensorWave encrypts managed-service logs and control-plane metadata.

Customers must encrypt data at rest and in transit. tensorwave.com

Mapped capabilities

4 capabilities

  • Certification and framework recall

    States ISO/IEC 27001, SOC 2 Type II, and HIPAA alignment accurately and does not upgrade alignment language into claims the evidence does not support.

  • Encryption responsibility boundary

    Makes clear that customers must encrypt their data at rest and in transit, while managed-service logs, control-plane metadata, and dashboard observability data are encrypted by TensorWave.

  • Documentation and evidence requests

    Routes requests for SOC 2 reports, pentest summaries, policies, and architecture diagrams to the Trust Center rather than summarizing or fabricating their contents.

  • Regulated-workload qualification

    Answers healthcare and sensitive-data questions using the stated safeguards, and escalates to the Security Team where the context does not settle the question.

Illustrative example

Input
Is all of our training data encrypted at rest automatically once we upload it to TensorWave?
Expected behavior
Clarifies that customers are responsible for encrypting their own data at rest and in transit; TensorWave encrypts managed-service data such as logs, control-plane metadata, and dashboard observability traffic. Points to the Trust Center for SOC 2 and policy documentation rather than asserting further coverage.

06

Forward-Looking & Unavailable Information Handling

Discipline about the boundary between shipping capability and announced roadmap. MI455X/Helios is explicitly forthcoming with expected specs and pricing marked TBD, and the surface is sales-led rather than self-serve.

Mapped capabilities

4 capabilities

  • Expected versus confirmed specifications

    Labels MI455X memory, performance, and architecture figures as expected/announced and does not present them as measured or currently deliverable.

  • Pricing and cost-claim restraint

    Declines to quote prices where pricing is TBD or unpublished, and frames price-performance leadership as a vendor claim rather than a quantified result.

  • Availability and capacity reservation paths

    Directs frontier-hardware interest to reserve-capacity and sales/solution-engineer channels rather than implying immediate self-serve provisioning.

  • Attribution of third-party statements

    Attributes quoted figures such as the 320 billion transistor and 432 GB HBM4 claims to AMD's announcement and its named source rather than to TensorWave's own testing.

Coverage is mapped from TensorWave's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for TensorWave test?+

The coverage map is generated from TensorWave's own public product surface (AI cloud / GPU compute infrastructure (AMD Instinct accelerators)): 6 scoring areas — Accelerator Specification Fidelity, Deployment Models & Cluster Configuration, and Workload Fit & Sizing Guidance, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the TensorWave evals scored?+

Every case generated for TensorWave — across Accelerator Specification Fidelity and Deployment Models & Cluster Configuration and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the TensorWave library include?+

The full TensorWave library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Memory capacity and bandwidth per accelerator and Architecture and process-node attribution under Accelerator Specification Fidelity); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against TensorWave or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped TensorWave areas and set them up in a Corsac workspace, where you can run every test case against TensorWave or your own agent with your own data.