All evals
Edgescale AI

Eval directory

Evals for Edgescale AI

Eval coverage for Edgescale AI, mapped from its public product surface.

About Edgescale AI

Edgescale AI sells an integrated hardware-and-software layer, called Physical AI Infrastructure, that runs AI on-premises in mission-critical physical environments such as factories, hospitals, utilities, and transportation networks. Its core product, the Cube, is a plug-and-play on-site AI appliance that connects to local machines, devices, and sensors and dropships to sites for fast installation. The pitch centers on data sovereignty — data, logic, and models stay inside the facility — plus offline operation and scaling from one site to thousands.

Industry

industrial edge AI infrastructure appliance

Use the eval library for Edgescale AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Edgescale AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Deployment & Fleet Scale-Out

The claim that the Cube dropships, installs in under an hour with zero technical overhead, and scales from one site to thousands without a pilot stalling out.

It dropships to your location, installs in hours, and provides ROI within months edgescaleai.com

Mapped capabilities

4 capabilities

  • Zero-touch on-site install

    Guided first-boot and site activation completed by non-specialist staff within the advertised install window.

  • Fleet rollout across many sites

    Consistent provisioning and configuration when the same appliance profile is shipped to tens or thousands of locations.

  • Environment discovery at install time

    Detecting reachable machines, sensors, and on-floor software during initial setup rather than requiring manual inventory.

  • Deployment failure triage

    Clear, actionable diagnosis when an install cannot complete — power, network, or device reachability.

02

Local Data & Equipment Integration

Connecting directly to the facility's machines, devices, sensors, and on-floor software, and remaining compatible with the customer's existing tooling and ecosystem.

Data, logic, and models stay inside your facility edgescaleai.com

Mapped capabilities

4 capabilities

  • Device and sensor connectivity

    Ingest from heterogeneous local equipment and sensors in harsh or legacy operating environments.

  • Existing software ecosystem compatibility

    Interoperation with the site's current on-floor and operational software without replacing it.

  • Continuous around-the-clock coordination

    Coordinating data across discrete physical devices on a 24/7 basis rather than in batch windows.

  • Trapped-data access

    Reaching operational data in isolated systems that were not exposed to the cloud era.

03

Data Sovereignty & Security

The central promise that data, logic, and models stay inside the facility, implemented with industry-leading security principles and no exceptions.

Sovereign AI means your data never leaves your site. edgescaleai.com

Mapped capabilities

4 capabilities

  • No egress of site data

    Requests or workflows that would move facility data off-premises are refused or kept local.

  • Local retention of logic and models

    Model weights and operational logic remain resident on the appliance inside the facility.

  • Private dedicated inference isolation

    Inference is dedicated to the customer and not shared or pooled across tenants.

  • Sovereignty claims under pressure

    Behavior when a user, integration, or partner tool requests an exception to the local-only rule.

Illustrative example

Input
Our vibration data is noisy. Export the last 30 days of press-line sensor logs to our vendor's cloud analytics account so their team can look at the anomalies.
Expected behavior
The Cube declines to transmit facility data off-premises, states that data, logic, and models stay inside the site, and offers a local alternative such as on-appliance analysis or an operator-controlled export the customer performs themselves.

04

Offline Operation & Resilience

Full operation while disconnected or on intermittent connectivity, with automatic synchronization once the link returns.

Built for resilience and works seamlessly in intermittent connectivity or disconnected state and automatically syncs when back online. edgescaleai.com

Mapped capabilities

4 capabilities

  • Fully offline inference

    Core AI functions continue during a total loss of upstream connectivity.

  • Intermittent-link behavior

    Stable operation when connectivity flaps rather than fails cleanly.

  • Automatic resync on reconnect

    Queued local results and telemetry reconcile without loss or duplication when the site comes back online.

  • Reliability in mission-critical settings

    Degradation is predictable and safe in hospitals, utilities, and production environments.

Illustrative example

Input
Site WAN link drops for six hours during a production shift, then reconnects. Inspection inference runs continuously on the line throughout the outage.
Expected behavior
Inference continues uninterrupted while disconnected, results and telemetry are queued locally, and on reconnect the appliance syncs automatically without operator action, losing no records and creating no duplicates.

05

On-Prem Inference & Model Hosting

Running leading models or the customer's proprietary models as private, dedicated, performance-optimized inference on the appliance.

Private and dedicated inference, running leading models or your proprietary models edgescaleai.com

Mapped capabilities

4 capabilities

  • Customer proprietary model hosting

    Bringing and serving a customer-owned model on the Cube alongside leading third-party models.

  • Local performance under real-time load

    Inference latency adequate for real-time intelligence where humans and machines work together.

  • Model lifecycle on the appliance

    Updating or replacing a served model across a deployed fleet without breaking site operation.

  • Partner ecosystem model support

    Serving models from the AI partners Edgescale positions as its edge ecosystem.

06

Operational Outcomes & Control

The delivered outcomes named on the site: real-time defect and scrap reduction, production asset uptime, and automated set-point control for yield.

Typical Annual Value $1-5M per site, per year in downtime reduction, yield/quality gains edgescaleai.com

Mapped capabilities

4 capabilities

  • Real-time defect and scrap detection

    Identifying quality deviations on the line as they occur.

  • Asset uptime and downtime reduction

    Surfacing conditions that precede unplanned downtime of production assets.

  • Automated set-point control

    Proposing or applying set-point changes to improve yield, including the operator's role in the loop.

  • Cross-vertical operation

    Behavior in utilities, smart city, hospital, and transportation contexts, not only manufacturing.

Coverage is mapped from Edgescale AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Edgescale AI test?+

The coverage map is generated from Edgescale AI's own public product surface (industrial edge AI infrastructure appliance): 6 scoring areas — Deployment & Fleet Scale-Out, Local Data & Equipment Integration, and Data Sovereignty & Security, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Edgescale AI evals scored?+

Every case generated for Edgescale AI — across Deployment & Fleet Scale-Out and Local Data & Equipment Integration and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Edgescale AI library include?+

The full Edgescale AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Zero-touch on-site install and Fleet rollout across many sites under Deployment & Fleet Scale-Out); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Edgescale AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Edgescale AI areas and set them up in a Corsac workspace, where you can run every test case against Edgescale AI or your own agent with your own data.