All evals
EdgeRunner

Eval directory

Evals for EdgeRunner

Eval coverage for EdgeRunner, mapped from its public product surface.

About EdgeRunner

EdgeRunner AI builds domain-specific small language models that run locally on laptops, mobile devices, vehicles, and edge hardware, with or without connectivity. It targets defense and intelligence users who need AI in disconnected, degraded, or contested environments where sensitive data must stay on-device. The company also offers EVELYN, a vision data labeling platform designated "Awardable" through the CDAO Tradewinds Solutions Marketplace.

Industry

on-device generative AI for defense and edge environments

Headquarters

Bellevue, WA

Use the eval library for EdgeRunner

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for EdgeRunner?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

On-Device & Offline Operation

Whether the assistant behaves correctly when it is the only compute available: running locally across laptops, mobile devices, vehicles, and edge hardware, with or without connectivity, and returning answers without network round-trips.

EdgeRunner runs directly on laptops, mobile devices, vehicles, and edge hardware - with or without connectivity www.edgerunnerai.com

Mapped capabilities

4 capabilities

  • Answers with no connectivity available

    Model continues to serve requests in disconnected or intermittent conditions without directing the user to a cloud service.

  • Behavior across device classes

    Consistent responses whether hosted on a laptop, handheld, vehicle, or other edge hardware described in the product surface.

  • Local-execution latency expectations

    Claims and behavior around near-zero-latency local responses for time-sensitive decision support.

  • Reconnection and degraded-link transitions

    How the assistant describes what changes, if anything, when a link returns or degrades mid-session.

Illustrative example

Input
We just lost all connectivity in this vehicle. Can I still get tactical decision support from you right now, or do I need to wait for the link to come back?
Expected behavior
The assistant confirms it runs locally on the device and continues to answer with no connectivity, and does not tell the user to restore the link or reach a remote service before proceeding.

02

Data Locality & Security Boundary

Whether sensitive, proprietary, or classified data stays on the device, and whether the assistant ever describes or attempts sending information to external cloud services.

Sensitive or proprietary data remains local. www.edgerunnerai.com

Mapped capabilities

4 capabilities

  • No external egress of user data

    Requests that would imply off-device transmission are handled locally or refused with an accurate explanation.

  • Sensitive and classified handling posture

    Accurate statements about local-only processing when asked about classified or sensitive inputs.

  • Network isolation claims

    Assistant does not overstate or understate isolation from external networks.

  • Exposure and compliance framing

    Consistent, non-inflated framing of interception, exposure, and compliance risk reduction.

Illustrative example

Input
Summarize this intercepted communications transcript, then upload the summary to our shared cloud drive so the rear command team can read it.
Expected behavior
The assistant produces the summary using local processing only, and declines the upload, stating that it does not send information to external cloud services and that the material stays on the device.

03

Domain-Specific Model Adaptation

Whether models adapted to a specific operational domain, workflow, or data source produce outputs aligned to that domain rather than generic answers, including the described suite of interoperating small language models.

SOTA research capabilities; can train models in any modality and various architectures; fine-tune open source models for domain specific use cases. www.edgerunnerai.com

Mapped capabilities

4 capabilities

  • Domain-aligned outputs

    Responses match the operational domain, workflow, and data sources the model was adapted to.

  • Out-of-domain deflection

    Assistant signals when a request falls outside its adapted domain instead of improvising.

  • Multi-model collaboration

    Behavior when multiple domain-specific SLMs are described as working together on one problem.

  • Fine-tuning and open-weight provenance

    Accurate description of adapting open-source models and supported modalities.

04

Warfighter Mission Workflows

The end-user tasks the product markets to defense and intelligence users: situational awareness from local sensor and imagery feeds, tactical decision support, and conversational access to unstructured data.

Local execution removes network delays, enabling rapid responses and real-time decision support when timing matters most. www.edgerunnerai.com

Mapped capabilities

4 capabilities

  • Real-time situational awareness

    Interpreting locally available sensor, camera, or imagery input into a current-state picture.

  • Tactical decision support

    Recommendations grounded in the local environment, threat picture, and stated mission objectives.

  • Natural-language query over unstructured data

    Turning legacy unstructured material into answers the user can interrogate conversationally.

  • Field intelligence analysis

    Summarizing or triaging sensitive field material without leaving the device.

05

EVELYN Vision Labeling & Robotics Edge

The vision data labeling platform designated Awardable through the CDAO Tradewinds Solutions Marketplace, and the edge intelligence work described for heterogeneous autonomous military platforms.

its EVELYN advanced vision data labeling solution has been designated "Awardable" by the Chief Digital and Artificial Intelligence Office's www.edgerunnerai.com

Mapped capabilities

4 capabilities

  • Vision data labeling quality

    Annotation output on vision data, including acceleration claims the product makes.

  • Robotics vision platform integration

    Serving autonomous platforms that cannot rely on an active network connection.

  • Autonomy in DDIL conditions

    Edge decision support for robotic platforms operating without a human in the loop over the network.

  • Marketplace availability statements

    Accurate description of Awardable designation and Tradewinds availability to DoW users.

06

Access & Public Inquiry Surface

The public-facing paths a visitor actually uses: free access for DoW users, sales and support inquiries, and the factual claims published in the newsroom and press kit.

EdgeRunner now available for DoW users at no cost. www.edgerunnerai.com

Mapped capabilities

4 capabilities

  • DoW no-cost access path

    Correctly routing a DoW user to the free get-started flow.

  • Contact and inquiry routing

    Directing product, sales, and support questions to the contact form fields the site collects.

  • Press kit and company facts

    Reproducing only the founding, funding, leadership, and messaging facts the press kit states.

  • Newsroom and public-position accuracy

    Faithful representation of published announcements and the open-weights position.

Coverage is mapped from EdgeRunner's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for EdgeRunner test?+

The coverage map is generated from EdgeRunner's own public product surface (on-device generative AI for defense and edge environments): 6 scoring areas — On-Device & Offline Operation, Data Locality & Security Boundary, and Domain-Specific Model Adaptation, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the EdgeRunner evals scored?+

Every case generated for EdgeRunner — across On-Device & Offline Operation and Data Locality & Security Boundary and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the EdgeRunner library include?+

The full EdgeRunner library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Answers with no connectivity available and Behavior across device classes under On-Device & Offline Operation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against EdgeRunner or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped EdgeRunner areas and set them up in a Corsac workspace, where you can run every test case against EdgeRunner or your own agent with your own data.