All evals
LA

Eval directory

Evals for Loman AI

Eval coverage for Loman AI, mapped from its public product surface.

About Loman AI

Loman AI is a 24/7 voice AI phone answering platform built for restaurants. It takes pickup and delivery orders, books reservations, answers customer FAQs, and takes card payments over the phone, syncing tickets and payments into POS and reservation systems like Square, Toast, SpotOn, and Clover. It is sold in Starter and Premium tiers (plus enterprise plans) to single-location, multi-unit, and enterprise brands.

Industry

restaurant voice AI phone agent

Headquarters

Austin, TX

Website

loman.ai

Use the eval library for Loman AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Loman AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Pickup & Delivery Order Taking

Capturing a takeout or delivery order over the phone accurately — items, modifiers, quantities, guest and fulfillment details — and confirming it back before it goes to the kitchen. This is the area Loman markets most heavily (order accuracy, out-of-stock menu updates, upsell), so it carries the most weight for an operator's buy decision.

99.6% order accuracy (industry best) loman.ai

Mapped capabilities

4 capabilities

  • Menu item and modifier capture

    Correct item, size, quantity, and modifier resolution from natural caller speech, including ambiguous or partial item names.

  • Order confirmation and correction

    Reading the order back and applying mid-call changes, removals, and substitutions without corrupting earlier items.

  • Out-of-stock and unavailable item handling

    Recognizing an unavailable item, saying so plainly, and offering a valid alternative rather than accepting the order.

  • Upsell and add-on behavior

    Offering relevant add-ons at appropriate moments without pressuring, repeating, or delaying order completion.

Illustrative example

Input
Caller: "Two large pepperoni pizzas for pickup. Actually, make the second one cheese instead, and add a side of ranch."
Expected behavior
The confirmed order contains exactly one large pepperoni pizza, one large cheese pizza, and one side of ranch, with fulfillment set to pickup. The agent reads the corrected order back before finalizing.

02

Phone Reservations

Booking, changing, and canceling reservations 24/7 by phone, including party size, date, time, and availability constraints. Reservation booking is a Premium-tier capability with named reservation-system sync, so correctness here is directly commercial.

Mapped capabilities

4 capabilities

  • Booking a new reservation

    Collecting party size, date, time, and guest name, and confirming the booked slot back to the caller.

  • Availability and alternative offers

    Handling requests for times that are unavailable by offering nearby valid options instead of over-booking.

  • Modify and cancel flows

    Locating an existing booking from caller-supplied details and applying a change or cancellation.

  • Large-party and special-request routing

    Recognizing requests that exceed normal booking rules and escalating or transferring instead of guessing.

03

Restaurant Knowledge & FAQ Accuracy

Answering questions about hours, location, menu contents, prices, and policies from the restaurant's own configured knowledge — and declining to answer what it does not know. Loman claims answers with 'total accuracy' from a restaurant knowledge graph, making grounding and abstention the core things to test.

Loman Answers 100% of calls so you never miss an order loman.ai

Mapped capabilities

4 capabilities

  • Hours, location, and logistics questions

    Answering from configured business info, including edge cases like holiday hours or 'are you open now'.

  • Menu, price, and dietary questions

    Answering item, price, and ingredient questions from the current menu without inventing details.

  • Abstention on unknown information

    Saying it does not have the information and offering a next step rather than fabricating an answer.

  • Allergen and dietary-restriction caution

    Handling allergen questions with appropriate caution and escalation instead of an unqualified safety assurance.

Illustrative example

Input
Caller: "Do you have a private room I can book for a work dinner, and what's the room fee?"
Expected behavior
With no private-dining information configured, the agent states it does not have that detail and offers to transfer or take a callback, rather than asserting a room exists or quoting any fee.

04

Phone Payments

Taking a credit card over the phone to close out an order, including how card data is handled conversationally and what happens when a payment fails. A Premium-tier capability where both correctness and caller-facing handling of sensitive data matter.

Loman securely takes credit card payments over the phone. loman.ai

Mapped capabilities

4 capabilities

  • Card capture and confirmation

    Collecting card details, confirming the charge amount, and confirming completion to the caller.

  • Sensitive-data handling in conversation

    Not reading back or unnecessarily repeating full card data, and not exposing it in transcripts or summaries.

  • Declined or failed payment recovery

    Explaining the failure, offering retry or pay-at-pickup, and leaving the order in a coherent state.

  • Total, tax, and tip accuracy

    Quoting an amount that matches the assembled order before authorizing the charge.

05

POS & Reservation System Sync

Getting the completed ticket, payment, and booking into the operator's existing stack — Square, Toast, SpotOn, Clover and reservation systems — and keeping menu and hours data current in the other direction. Integration depth is Loman's stated differentiator and the most common source of silent operator-visible failure.

Loman syncs with top POS systems like Square, Toast, SpotOn and Clover to streamline orders and reservations. loman.ai

Mapped capabilities

4 capabilities

  • Ticket fidelity into the POS

    Items, modifiers, fulfillment type, and guest details arriving in the POS matching what was confirmed on the call.

  • Reservation write-through

    Bookings, changes, and cancellations reflected in the connected reservation system.

  • Menu and hours ingestion

    Automatic menu and business-info updates flowing into what the agent tells callers.

  • Sync failure visibility

    Surfacing a failed or partial push to the operator instead of dropping the order silently.

06

Live Call Control & Recovery

How the agent behaves as a phone system under real conditions: many simultaneous calls, interruptions, unclear audio, callers who want a human, and follow-up SMS. Loman markets multi-call handling, human-level interruption handling, and automated transfer, so these are load-bearing claims rather than incidental behavior.

Mapped capabilities

4 capabilities

  • Interruption and barge-in handling

    Yielding when the caller talks over the agent and resuming from the correct point in the flow.

  • Transfer and escalation to staff

    Recognizing complaints, refunds, or explicit human requests and transferring with context instead of looping.

  • Unclear input and repair

    Handling misheard names, addresses, and numbers by re-asking narrowly rather than guessing or restarting.

  • Outbound SMS and call artifacts

    Confirmation and status SMS content matching the call outcome, and transcripts reflecting what happened.

Coverage is mapped from Loman AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Loman AI test?+

The coverage map is generated from Loman AI's own public product surface (restaurant voice AI phone agent): 6 scoring areas — Pickup & Delivery Order Taking, Phone Reservations, and Restaurant Knowledge & FAQ Accuracy, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Loman AI evals scored?+

Every case generated for Loman AI — across Pickup & Delivery Order Taking and Phone Reservations and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Loman AI library include?+

The full Loman AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Menu item and modifier capture and Order confirmation and correction under Pickup & Delivery Order Taking); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Loman AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Loman AI areas and set them up in a Corsac workspace, where you can run every test case against Loman AI or your own agent with your own data.