All evals
VA

Eval directory

Evals for VESSL AI

Eval coverage for VESSL AI, mapped from its public product surface.

About VESSL AI

VESSL AI is a GPU cloud and orchestration layer that gives AI teams on-demand and reserved NVIDIA GPUs (A100 through B300/GB300) at published hourly rates. Its VESSL Cloud product offers persistent single-node GPU workspaces, one-command jobs, and multi-node clusters, with SSH access from your own editor and preconfigured container environments. A CLI (vesslctl) with MCP integration lets users drive the platform from the terminal.

Industry

GPU cloud / AI infrastructure platform

Website

vessl.ai

Use the eval library for VESSL AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for VESSL AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

GPU catalog and hardware specs

Answering questions about the published accelerator lineup — model, architecture, memory, bandwidth, and interconnect — without drifting from the stated numbers or inventing SKUs.

Mapped capabilities

4 capabilities

  • Model-to-architecture and memory lookup

    B300/B200/GB300 (Blackwell), H200/H100 (Hopper), A100 (Ampere), L40S (Ada Lovelace) with their stated VRAM.

  • Bandwidth and throughput figures

    HBM3e/HBM3/HBM2e/GDDR6 bandwidth and the published FP8/BF16 throughput per part.

  • Multi-node interconnect availability

    Which parts are listed with 400 Gbps NDR InfiniBand for multi-node versus single-node only.

  • Workload fit guidance

    Repeating the site's own positioning (fine-tuning value pick, large-scale training workhorse, image/video generation) without overstating it.

02

Pricing, billing, and cost estimation

Quoting published hourly rates, applying them to a stated GPU count and usage window, and handling reserved discounts and storage tiers accurately.

Hourly rates published to the cent, from the A100 to the B300. vessl.ai

Mapped capabilities

4 capabilities

  • Published hourly rate recall

    B300 $7.38, B200 $5.88, H200 $4.38, H100 $2.98, L40S $1.80, A100 $1.48 per hour; GB300 reserved-only with no listed rate.

  • Monthly cost arithmetic

    Rate × GPU count × hours for the 730h (24/7) and 176h (business hours) presets used by the estimator.

  • Reserved discount and billing model

    Per-second billing, up to 15% off reserved, and reserved terms quoted with sales rather than published.

  • Storage tier distinction

    Warm cluster storage for active datasets and checkpoints versus cold object storage for archives.

Illustrative example

Input
What would a single A100 cost me if I only run it during business hours, about 176 hours a month?
Expected behavior
States the published A100 SXM rate of $1.48 per hour and computes roughly $260 for 176 hours, notes billing is per-second, and frames the figure as an estimate at published rates rather than a final quote.

03

Access policy and sales routing

Distinguishing what a user can start on their own from what requires a sales conversation, and routing accordingly instead of promising unavailable self-serve paths.

H200 and up run as reserved capacity; final terms are quoted with our team. vessl.ai

Mapped capabilities

4 capabilities

  • Self-serve versus on-request parts

    H100, A100, and L40S self-serve; B300, B200, GB300, and H200 listed as on request.

  • Reserved-only capacity

    GB300 offered as reserved only, sized 288GB × N, positioned for cluster-scale builds.

  • Estimate-versus-quote framing

    Presenting calculator output as an estimate at published rates and deferring final terms to sales.

  • Compliance and trust claims

    SOC 2 Type II as stated on the pricing page, without extending it to unstated certifications.

Illustrative example

Input
Can I put in a credit card and spin up an H200 right now, the way I would an H100?
Expected behavior
Explains that H200 is listed as on request rather than self-serve and routes the user to sales, while noting H100, A100, and L40S are the self-serve options. May cite the published H200 rate of $4.38 per hour.

04

Compute workflow surfaces

Explaining and selecting among the three VESSL Cloud execution modes — workspace, job, and multi-node cluster — for a described task.

Persistent GPU instances, one-command jobs, and the same environment from first experiment to production. vessl.ai

Mapped capabilities

4 capabilities

  • Persistent GPU workspace

    Single node up to 8 GPUs (A100, H100, L40S), SSH into the user's own IDE, persistent environment, image, and volumes.

  • One-command jobs

    Submitting a training run, evaluation, or preprocessing pipeline to run to completion in the same environment.

  • Scaling to multi-node clusters

    When one node is not enough, and which parts carry the multi-node InfiniBand listing.

  • Preconfigured container environments

    Container-based images that the site credits with removing OS and driver conflicts.

05

Terminal and agentic control

Driving the platform from the CLI and agent surfaces described in the product blog and cloud page, without fabricating commands or capabilities.

Mapped capabilities

4 capabilities

  • vesslctl workflow coverage

    The official VESSL Cloud CLI as the terminal path for the described workflow.

  • MCP integration and bundled Claude skill

    Native MCP integration and the bundled skill shipped with the CLI.

  • Agent-driven availability and launch

    The demonstrated flow of checking GPU availability, launching a job on an 8×H100 node, and verifying it.

  • Refusing to invent CLI surface

    Declining to supply flags, subcommands, or endpoints not evidenced in the supplied material.

06

Company grounding and localization

Handling company facts, published content, and the English/Korean site pair with fidelity to what is actually stated.

Mapped capabilities

4 capabilities

  • Company facts and milestones

    Founded 2020, $16M+ total funding, 100+ customers, 3.4x YoY revenue growth, VESSL Cloud launched February 2026.

  • Named partnerships and recognition

    Hyundai Motor (autonomous driving), Tmap Mobility (AI agents), CB Insights' 2025 AI Agent Tech Stack listing.

  • Blog and tutorial content recall

    Published pieces such as the A100/H100/B200 LoRA cost benchmark and the Gemma 4 fine-tuning tutorial.

  • English/Korean parity

    Serving the same catalog, rates, and CTAs consistently across the /en and /ko surfaces.

Coverage is mapped from VESSL AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for VESSL AI test?+

The coverage map is generated from VESSL AI's own public product surface (GPU cloud / AI infrastructure platform): 6 scoring areas — GPU catalog and hardware specs, Pricing, billing, and cost estimation, and Access policy and sales routing, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the VESSL AI evals scored?+

Every case generated for VESSL AI — across GPU catalog and hardware specs and Pricing, billing, and cost estimation and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the VESSL AI library include?+

The full VESSL AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Model-to-architecture and memory lookup and Bandwidth and throughput figures under GPU catalog and hardware specs); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against VESSL AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped VESSL AI areas and set them up in a Corsac workspace, where you can run every test case against VESSL AI or your own agent with your own data.