All evals
Fleet AI

Eval directory

Evals for Fleet AI

Eval coverage for Fleet AI, mapped from its public product surface.

About Fleet AI

Fleet AI builds simulated worlds and real-world challenges used to study and shape how artificial intelligences behave. The public site is largely a brand, careers, and legal presence, with a members-only dashboard and a "Notes" section about the worlds the team builds. Beyond the positioning statement, the pages disclose little concrete product functionality, pricing, or customer-facing capability detail.

Industry

AI agent evaluation and simulation environments

Headquarters

San Francisco (SOMA), with a satellite office in New York City (Chelsea); entity registered in Delray Beach, Florida

Use the eval library for Fleet AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Fleet AI?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Company positioning and identity

The single positioning statement — Fleet creates simulated worlds and real-world challenges to better understand and shape the behavior of artificial intelligences — plus the 'fleet' definition gloss that recurs sitewide. The main risk is over-reading this brand copy into concrete product claims.

creates simulated worlds and real-world challenges to better understand and shape the behavior of artificial intelligences www.fleetai.com

Mapped capabilities

4 capabilities

  • Positioning statement fidelity

    Restates the mission line without adding capabilities, customers, or outcomes not present on the site.

  • 'Fleet' term gloss

    Handles the noun/verb definition and the Fleet Street reference as brand copy, not product taxonomy.

  • Product-detail absence

    Says plainly that pricing, features, and customer-facing capability detail are not disclosed publicly.

  • Sitewide navigation inventory

    Accounts for the actual set of surfaces: Dashboard, Careers, Notes, About, Brand, and the legal pages.

02

Brand assets and naming

The /brand page is the most functionally concrete public surface: it fixes the company's name of record and offers three downloadable SVG asset variants. Naming and asset-variant confusion are the realistic failure modes.

Mapped capabilities

4 capabilities

  • Name of record

    Uses 'Fleet AI' as instructed on the brand page, including in prose where 'Fleet' alone is ambiguous.

  • Asset variant selection

    Distinguishes full logo, emblem, and wordmark, and picks the right one for a stated placement.

  • Asset format and delivery

    Reports that assets are offered as SVG downloads and does not promise other formats or color variants.

  • Usage guidance boundaries

    Declines to invent clearspace, color, or co-branding rules the page does not state.

Illustrative example

Input
We're adding Fleet to our partner page. What should we call them, and which logo file fits a 32px-tall header slot?
Expected behavior
Uses the name Fleet AI as the brand page instructs, and recommends the emblem for a constrained 32px header while noting the full logo and wordmark are the other available SVG downloads. It does not invent clearspace, color, or minimum-size rules the page never states.

03

Careers and hiring funnel

The most detailed public page: work-mode options, a named benefits set, office locations, and an open-positions area with an explicit fallback path for candidates who see no perfect fit. Candidate-facing accuracy matters most here.

Mapped capabilities

4 capabilities

  • Work mode and locations

    Covers in-person (SF, NYC) versus remote, and the SOMA HQ / Chelsea satellite detail.

  • Benefits enumeration

    Reproduces equity, health/vision/dental, meal stipend, gym access, unlimited PTO, 401k, and hardware without embellishment.

  • Open roles handling

    Avoids asserting specific titles, levels, or compensation that the supplied page text does not list.

  • No-fit contact path

    Routes a candidate without a matching role to the direct-contact fallback the page offers.

04

Notes and editorial content

A team-authored section titled 'Learning(s) from the Worlds we Build,' with All/Team filtering and sorting, and at least one dated team profile. Risk is treating a personnel story as a product or research disclosure.

Learning(s) from the Worlds we Build www.fleetai.com

Mapped capabilities

3 capabilities

  • Note retrieval and attribution

    Attributes the Feb 11, 2026 note to Fred Havemeyer, Founding Member of Technical Staff, with its customer-to-first-engineer framing.

  • Filter and sort semantics

    Explains the All/Team filter and sort affordances without inventing categories.

  • Editorial vs. product claims

    Keeps note content as team narrative and does not convert it into capability or roadmap claims.

06

Cross-surface consistency and access boundary

The public pages carry an unresolved split between the fleetai.com site and the usefleet.ai domain named in the legal text, plus a fleet.so contact address; separately, Dashboard sits behind a members-only boundary with a sitewide status indicator.

Mapped capabilities

4 capabilities

  • Domain and entity discrepancy

    Surfaces the fleetai.com / usefleet.ai / fleet.so mismatch as an observed inconsistency rather than silently picking one.

  • Policy currency

    Notes that the legal pages are dated 9/16/2023 against a 2026 site copyright, without asserting which is stale.

  • Gated dashboard boundary

    Treats Dashboard as members-only and does not describe or infer post-login functionality.

  • Status indicator scope

    Reads 'All Systems Operational' as a sitewide status signal, not evidence of a specific product's uptime.

Illustrative example

Input
Your terms say the site is usefleet.ai but I'm reading this on fleetai.com. Which one is the real company site?
Expected behavior
Confirms both strings appear as published — the legal pages name usefleet.ai while the live pages are served from fleetai.com — flags the discrepancy explicitly, and points to the listed contact for authoritative resolution instead of declaring one domain correct.

Coverage is mapped from Fleet AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Fleet AI test?+

The coverage map is generated from Fleet AI's own public product surface (AI agent evaluation and simulation environments): 6 scoring areas — Company positioning and identity, Brand assets and naming, and Careers and hiring funnel, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Fleet AI evals scored?+

Every case generated for Fleet AI — across Company positioning and identity and Brand assets and naming and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Fleet AI library include?+

The full Fleet AI library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Positioning statement fidelity and 'Fleet' term gloss under Company positioning and identity); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Fleet AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Fleet AI areas and set them up in a Corsac workspace, where you can run every test case against Fleet AI or your own agent with your own data.