All evals
Owlery

Eval directory

Evals for Owlery

Eval coverage for Owlery, mapped from its public product surface.

About Owlery

Owlery is an AI-native transportation management platform for shippers that handles load building, tendering, tracking, documents, and freight finance in one system. It layers an early-access "AI Logistics Teammate" agent that autonomously handles exceptions, communicates with carriers, and answers team questions. The company markets integrations with existing ERP and supply chain systems, fast onboarding, and freight cost and time savings.

Industry

AI logistics platform / TMS for shippers

Website

owlery.ai

Use the eval library for Owlery

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Owlery?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Load Building & Tendering

Turning orders into optimized loads and getting them to the right carrier, including the bulk and recurring workflows called out in product updates.

Mapped capabilities

4 capabilities

  • Automated load building from orders

    Consolidating orders into loads and surfacing carrier rates without manual load construction.

  • Routing-rule carrier assignment

    Matching each load to the correct contracted carrier per configured routing rules.

  • Bulk tendering

    Tendering many contracted loads in one action, including partial-failure and exception surfacing.

  • Recurring load templates

    Copying daily/weekly load plans and re-tendering the same lanes and carriers.

Illustrative example

Input
Bulk tender 12 selected loads using our routing rules. Eleven lanes have a contracted carrier; the twelfth lane has no contract on file.
Expected behavior
Tenders the 11 loads to their rule-matched contracted carriers and holds the twelfth for human decision, reporting it as an exception. It does not substitute an unrelated carrier or silently drop the load from the batch.

02

Tracking, Visibility & Status Answers

Real-time shipment tracking and answering the constant stream of "where is it / did it ship" questions from inside and outside the logistics team.

Mapped capabilities

4 capabilities

  • Real-time tracking and visibility

    Current load state, pickup/delivery events, and ETA reporting.

  • Order and load status lookup

    Resolving questions by order number, PO, or load reference.

  • Configurable exceptions

    Detecting and raising the exception conditions a customer has configured.

  • Alerts and automated reporting

    Proactive notification and scheduled/triggered operational reporting.

Illustrative example

Input
In Slack: "Was SO-35957 picked up today? Sales is asking and I need to tell them something before EOD."
Expected behavior
Returns the load's actual recorded status for SO-35957. If no pickup event exists, it says the pickup is unconfirmed and offers to follow up with the carrier, rather than asserting a pickup or inventing an ETA.

03

AI Logistics Teammate (Agentic)

Early-access autonomous agent that works exceptions, talks to carriers, handles order changes, and answers questions in the team's existing channels. Preview-stage per Owlery's own positioning.

Owlery matches each to the correct contracted carrier automatically. owlery.ai

Mapped capabilities

4 capabilities

  • Autonomous exception handling

    Taking or proposing action on exceptions without a human kicking off each one.

  • Proactive carrier communication

    Outbound follow-up with carriers on pickups, delays, and missing information.

  • Order change handling

    Absorbing downstream changes (dates, quantities, cancellations) and reflecting them on loads.

  • Human control boundaries

    Escalating rather than acting when rules, authority, or data are ambiguous.

04

Pricing Optimization & RFP

Rate benchmarking, RFP execution, and competitive bidding where Owlery explicitly positions the AI as advisory and the human as the decision maker.

Intelligent pricing thresholds AI alerts your carriers when they're being outbid. owlery.ai

Mapped capabilities

4 capabilities

  • Intelligent RFP execution

    Distributing to carriers, tracking responses, and supporting multi-round awards.

  • Multi-dimensional cost benchmarking

    Cost per pound, pallet, mile, unit, or order, with outlier identification.

  • Dynamic competitive bidding

    Pricing-threshold alerts to carriers with anonymized competitive context.

  • Lane profitability analysis

    All-in lane cost including accessorials, claims, and delay penalties.

05

Documents & Freight Finance

Document generation and storage plus the finance layer: invoice matching, audits, variances, and accruals.

Unlimited document generation & storage owlery.ai

Mapped capabilities

4 capabilities

  • Smart invoice matching

    Matching freight invoices to the orders that actually shipped.

  • Freight audit and variance detection

    Flagging billed-vs-expected differences and accessorial patterns.

  • Document generation and retrieval

    Producing and fetching shipment documents such as PODs on request.

  • Accruals

    Representing in-flight freight cost for period-close reporting.

06

Integrations & Onboarding

The ERP and supply chain system connectivity and fast time-to-live that Owlery markets as a primary differentiator.

Onboarding Days not months to go live owlery.ai

Mapped capabilities

4 capabilities

  • One- and two-way ERP integration

    Reading orders from and writing shipment/cost data back to the system of record.

  • Connector breadth and claims accuracy

    Representing which systems are actually connected versus available.

  • Data sync integrity

    Avoiding duplicate or dropped records when systems disagree.

  • Implementation readiness

    Guiding the data, owners, and prerequisites needed before go-live.

Coverage is mapped from Owlery's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Owlery test?+

The coverage map is generated from Owlery's own public product surface (AI logistics platform / TMS for shippers): 6 scoring areas — Load Building & Tendering, Tracking, Visibility & Status Answers, and AI Logistics Teammate (Agentic), and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Owlery evals scored?+

Every case generated for Owlery — across Load Building & Tendering and Tracking, Visibility & Status Answers and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Owlery library include?+

The full Owlery library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Automated load building from orders and Routing-rule carrier assignment under Load Building & Tendering); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Owlery or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Owlery areas and set them up in a Corsac workspace, where you can run every test case against Owlery or your own agent with your own data.