All evals
A

Eval directory

Evals for Armada

Eval coverage for Armada, mapped from its public product surface.

About Armada

Armada Edge Platform is a full-stack edge computing platform made up of four products: Atlas, Galleon, Bridge, and Marketplace. It pairs ruggedized, containerized modular data centers (Galleon, including the megawatt-scale Leviathan form factor) with software for connected-asset monitoring (Atlas) and GPU orchestration and monetization (Bridge). It targets remote and regulated operations — offshore rigs, defense missions, mining sites — that need fast deployment and data sovereignty.

Industry

edge AI compute infrastructure platform

Use the eval library for Armada

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Armada?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Platform composition and product routing

Correctly distinguishing the four AEP products and directing a stated need to the right one, without blending Atlas, Galleon, Bridge, and Marketplace into a single undifferentiated offering.

Mapped capabilities

4 capabilities

  • Four-product decomposition

    Names Atlas, Galleon, Bridge, and Marketplace and states each one's distinct role as described on the product pages.

  • Need-to-product routing

    Maps an operational need (asset visibility, on-site compute, GPU orchestration, partner AI software) to the correct product.

  • Atlas vs. Bridge boundary

    Keeps connected-asset monitoring (Atlas) separate from GPU orchestration and monetization (Bridge).

  • Galleon form-factor family

    Places Beacon, Scout, Triton, Cruiser, and Leviathan as configurations within the Galleon line rather than separate products.

02

Galleon and Leviathan infrastructure specifications

Quoting the published capacity, footprint, and timeline figures for modular data centers precisely, and declining to invent specs the pages do not state.

Leviathan delivers 1.77 MW of AI infrastructure, operational in 12 weeks. www.armada.ai

Mapped capabilities

4 capabilities

  • Leviathan capacity figures

    1.77 MW IT envelope, up to 576 NVIDIA GB300 GPUs per unit, 2.2 MW cooling, .12 acres per unit, 50 MW+ cluster path.

  • Deployment timeline claims

    Galleon in 60 days and Leviathan in 12 weeks, presented against the cited 24-month traditional data center industry average.

  • Integrated-system framing

    Describes Leviathan as a complete system including power, cooling, compute, networking, controls, and software — not a shell or enclosure.

  • Configuration sizing

    Matches a stated power need to a nearest-fit Galleon configuration and flags requests outside the practical range.

Illustrative example

Input
How much AI capacity does one Leviathan unit deliver, and how fast can we have it running?
Expected behavior
States 1.77 MW IT envelope and up to 576 NVIDIA GB300 GPUs per unit, operational in 12 weeks. Does not inflate the per-unit numbers toward the 50 MW+ cluster scale path or the 5–20 MW optimization range, and keeps those framed as cluster-level figures if mentioned.

03

Sovereignty, air-gap, and compliance constraints

Treating data residency, air-gapped operation, and regulatory compliance as binding deployment requirements that shape the recommendation, since these are the platform's stated reason for existing.

Bridge runs directly on your infrastructure, enabling fully data-sovereign AI cloud services www.armada.ai

Mapped capabilities

4 capabilities

  • Data residency boundaries

    Explains keeping data at the site, within the organization, or inside national borders, and who controls what leaves.

  • Air-gapped configurations

    Identifies Triton and Cruiser as configurable to operate fully air-gapped with no outside network exposure.

  • Bridge on-premises sovereignty

    Bridge runs on the operator's own infrastructure to meet regional compliance and security requirements.

  • Constraint-driven recommendation

    Lets a stated sovereignty or air-gap requirement rule out otherwise attractive connected options.

Illustrative example

Input
We're a defense customer. Our site cannot touch any outside network, ever. Which Galleon configuration should we look at, and does that change how Atlas works for us?
Expected behavior
Names Triton and Cruiser as the configurations Armada states can operate fully air-gapped with no outside network exposure, and treats the air-gap as a hard constraint on the recommendation. Notes that cloud-dependent monitoring behavior under full air-gap is not specified in available material rather than asserting it.

04

Atlas connected-asset and drone operations

Supporting the day-to-day monitoring workflow Atlas exposes: multi-vendor connectivity fleets, pooled data plans, and drone mission lifecycle from pre-flight through post-flight review.

Monitor connectivity, beam performance, plan usage, and terminal health for your Starlink and Viasat fleet in one unified console. www.armada.ai

Mapped capabilities

4 capabilities

  • Starlink and Viasat fleet monitoring

    Connectivity, beam performance, plan usage, and terminal health across both vendors in one console.

  • Pooled data plan mechanics

    Sharing data across terminals for optimized usage and transparent monitoring without overage fees.

  • Drone mission lifecycle

    Live telemetry and video, AI-flagged incidents along the flight path, and searchable replay by location, time, and asset.

  • Multi-vendor failover

    Explains multi-vendor connectivity as a resilience and failover posture across the asset fleet.

05

Bridge GPU orchestration, tenancy, and monetization

Reasoning about turning owned GPU capacity into a cloud-like service: isolation between tenants, elastic allocation, and revenue from spare capacity.

Mapped capabilities

4 capabilities

  • Hard multi-tenant isolation

    Describes hard isolation between tenants as the basis for safely sharing a cluster.

  • Elastic resource allocation

    Dynamic workload optimization and scaling across internal and external GPU resources.

  • GPU-as-a-Service provisioning

    Provisioning and managing GPUs across datacenter, cloud, and edge deployments from one control plane.

  • Capacity monetization

    Framing spare GPU capacity as a revenue-generating service rather than idle infrastructure.

06

Marketplace partner AI services

Handling the validated partner software layer that sits on Bridge, including which service categories are available and how they reach infrastructure.

Mapped capabilities

3 capabilities

  • Service category coverage

    Model serving, optimization, agentic AI, computer vision, data governance, and security.

  • Deployment onto Bridge

    Partner-validated software deploys directly on Bridge to turn infrastructure into AI capabilities.

  • Partner-validation framing

    Distinguishes validated partner offerings from arbitrary third-party software the operator sources independently.

Coverage is mapped from Armada's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Armada test?+

The coverage map is generated from Armada's own public product surface (edge AI compute infrastructure platform): 6 scoring areas — Platform composition and product routing, Galleon and Leviathan infrastructure specifications, and Sovereignty, air-gap, and compliance constraints, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Armada evals scored?+

Every case generated for Armada — across Platform composition and product routing and Galleon and Leviathan infrastructure specifications and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Armada library include?+

The full Armada library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Four-product decomposition and Need-to-product routing under Platform composition and product routing); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Armada or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Armada areas and set them up in a Corsac workspace, where you can run every test case against Armada or your own agent with your own data.