All evals
SambaNova

Eval directory

Evals for SambaNova

Eval coverage for SambaNova, mapped from its public product surface.

About SambaNova

SambaNova sells an AI inference stack built around its Reconfigurable Dataflow Unit (RDU) chip, currently the fifth-generation SN50. The product family spans the SambaRack air-cooled hardware rack, the SambaStack full-stack on-prem or hosted deployment, the SambaManaged turnkey inference cloud run in a customer's own data center, and SambaCloud, a developer API for fast inference on large open-source models. Positioning emphasizes speed and throughput on very large models, energy efficiency per token, and OpenAI-compatible endpoints for easy migration.

Industry

AI inference infrastructure (chips, racks, and cloud platform)

Use the eval library for SambaNova

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for SambaNova?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

SambaCloud Developer API & Migration

The developer-facing inference API: OpenAI-compatible endpoints, Anthropic Messages API support, key and base-URL configuration, prompt caching, and the ecosystem integrations used to get a first call working.

scaling up to 256 RDUs working together for inference on the SN50 sambanova.ai

Mapped capabilities

4 capabilities

  • OpenAI-compatible endpoint migration

    Swapping an existing OpenAI client over by setting the API key, base URL, and model id; what does and does not need to change.

  • Anthropic Messages API support

    Recognizing that SambaCloud also exposes an Anthropic-compatible surface and when a caller would choose it.

  • Prompt caching behavior

    Explaining prompt caching as a latency and cost lever on supported models without overstating guarantees.

  • Integrations and tooling

    Getting started through named integrations such as CrewAI, Hugging Face, Cline, and AWS.

Illustrative example

Input
I already use the OpenAI Python SDK. What is the smallest change to run gpt-oss-120b on SambaCloud instead of OpenAI?
Expected behavior
Explains that the OpenAI-compatible endpoint means keeping the same SDK: set the API key to a SambaNova key, point the base URL at SambaNova, and select the model. No client rewrite is required.

02

Model Catalog & Selection

Which large open-source models are served, what modalities they cover, and how a user picks among them for a given workload on a speed-versus-size tradeoff.

Mapped capabilities

4 capabilities

  • Supported model families

    Answering which open-source families are available, including DeepSeek, Llama, Qwen, Gemma, MiniMax, and gpt-oss variants.

  • Modality coverage

    Distinguishing text, image, and audio capability across catalog entries rather than assuming uniform support.

  • Model selection guidance

    Matching a described workload (agentic, coding, high-volume chat) to a catalog model using grounded attributes.

  • Benchmark and speed claims

    Attributing inference-speed claims to their stated third-party sources rather than presenting them as unsourced fact.

03

RDU & Dataflow Architecture Explanation

Technical explanation of the fifth-generation SN50 RDU, the Dataflow Architecture, and the memory design — the differentiation story a technical buyer will probe.

SambaRack SN50 delivers 5X more compute per accelerator and 4X more network bandwidth than the previous generation. sambanova.ai

Mapped capabilities

4 capabilities

  • Dataflow versus kernel-by-kernel execution

    Explaining PCU/PMU grids, the streaming pipeline, and why intermediate activations stay on-chip.

  • Data movement as the bottleneck

    Articulating why off-chip data movement, not raw compute, is framed as the central cost.

  • Three-tier memory and model switching

    How tiered memory supports very large models and fast switching between models.

  • SN50 versus SN40 generational deltas

    Stating the specific claimed gains (compute, network bandwidth, scale to 256 accelerators) without inflating them.

04

SambaRack Hardware & Sizing

The air-cooled rack itself: chip count, scale-out limits, model and context ceilings, power and floor-space planning, and the sizing walkthrough for multi-megawatt buildouts.

SambaStack is powered by SambaRack, the most efficient rack for AI using an average of 10 kW of power sambanova.ai

Mapped capabilities

4 capabilities

  • Rack composition and scale-out

    16 RDU chips per rack, scaling across racks and up to 256 accelerators over the interconnect.

  • Model and context ceilings

    Handling stated limits on parameter count and context length for a single rack.

  • Power and cooling envelope

    Air-cooled operation at roughly 10 kW per rack and its implications for existing data centers.

  • Deployment sizing walkthrough

    Guiding a user through energy, floor space, and cluster scale estimates for a planned deployment.

05

Deployment Models & Fit

Choosing correctly among SambaCloud, SambaStack, SambaManaged, and SambaRack — the disambiguation surface where a wrong recommendation is expensive.

a single SambaRack SN50 can run models of up to 10 trillion parameters sambanova.ai

Mapped capabilities

4 capabilities

  • Offering disambiguation

    Correctly separating the developer API, the full-stack deployment, the managed in-customer-DC cloud, and the underlying rack.

  • On-prem versus hosted SambaStack

    Representing both deployment options and the dedicated-infrastructure framing.

  • Turnkey timelines and operating model

    Deployment-in-90-days and no-specialized-AI-expertise-required claims, stated as vendor claims.

  • Model bundles and hot-swapping

    Pre-configured bundles per rack and swapping between bundles at inference time to serve more models per footprint.

Illustrative example

Input
We are a European bank with our own data center and no in-house AI team. We want inference for internal groups with data kept on-shore. Which SambaNova offering fits?
Expected behavior
Recommends SambaManaged, the fully managed inference cloud that runs in the customer's own data center, and distinguishes it from the public SambaCloud API. Cites grounded fit factors such as the 90-day launch, air-cooled racks at about 10 kW, and on-shore data residency.

06

Privacy, Sovereignty & Claim Discipline

Data handling and residency commitments, sovereign deployment positioning, and the discipline of not overstating unverified performance, valuation, or customer claims.

SambaCloud never sees or collects any of your data or user prompts, ensuring full data privacy. sambanova.ai

Mapped capabilities

4 capabilities

  • Data privacy posture

    Representing the stated position that SambaCloud does not see or collect customer data or prompts.

  • Sovereign and on-shore deployment

    Keeping data, models, and compliance in-region for named government and enterprise geographies.

  • Attribution of third-party benchmarks

    Citing Artificial Analysis and SemiAnalysis as the sources of speed claims rather than asserting them directly.

  • Corporate and customer claims

    Handling financing, valuation, and named-customer statements as time-sensitive claims requiring a source.

Coverage is mapped from SambaNova's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for SambaNova test?+

The coverage map is generated from SambaNova's own public product surface (AI inference infrastructure (chips, racks, and cloud platform)): 6 scoring areas — SambaCloud Developer API & Migration, Model Catalog & Selection, and RDU & Dataflow Architecture Explanation, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the SambaNova evals scored?+

Every case generated for SambaNova — across SambaCloud Developer API & Migration and Model Catalog & Selection and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the SambaNova library include?+

The full SambaNova library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, OpenAI-compatible endpoint migration and Anthropic Messages API support under SambaCloud Developer API & Migration); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against SambaNova or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped SambaNova areas and set them up in a Corsac workspace, where you can run every test case against SambaNova or your own agent with your own data.