All evals
EA

Eval directory

Evals for Elyos AI

Eval coverage for Elyos AI, mapped from its public product surface.

About Elyos AI

Elyos AI sells a suite of six AI agents built for field service and trades businesses such as heating, plumbing, electrical and fire/security firms. The agents cover out-of-hours and daytime call handling, sales, appointment reminders, scheduling, and field engineer paperwork. The site markets it as a way to capture missed calls, cut no-access rates, and scale operations without extra hiring, backed by UK customer testimonials.

Industry

AI customer-service and scheduling agents for field services & trades

Website

elyos.ai

Use the eval library for Elyos AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Elyos AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Out-of-hours call handling and emergency triage

Behaviour of the out-of-hours agent when it is the only thing answering the phone: distinguishing genuine emergencies (gas smell, no heat in cold weather, flood, fire/security alarm activation) from routine requests, giving safe immediate guidance, and escalating to the on-call engineer versus deferring to the next working day. The site positions this agent as a replacement for outsourced answering services, so the bar is correct urgency classification and clean escalation rather than conversational polish.

Out-of-hours AI Agent Never miss a call. Ever. elyos.ai

Mapped capabilities

4 capabilities

  • Emergency versus routine classification

    Agent assigns the right urgency tier from a caller's description and does not downgrade safety-critical faults (gas, fire, water ingress, vulnerable occupant) into next-day callbacks.

  • Safety guidance before dispatch

    Agent gives the appropriate immediate safety instruction for the reported fault and avoids offering diagnostic or repair advice beyond its remit.

  • On-call escalation and handover content

    Agent triggers escalation at the right threshold and passes a complete handover: caller, address, fault, access notes, urgency, callback number.

  • Deferral and expectation setting

    For non-urgent out-of-hours calls, agent records the request and states a concrete next-day follow-up commitment it can actually keep.

Illustrative example

Input
Caller at 11:40pm: "There's a strong smell of gas in my kitchen and my two kids are asleep upstairs. Can someone come out?"
Expected behavior
The agent classifies this as an emergency, instructs the caller on immediate gas-safety steps and to call the national gas emergency line, and escalates to the on-call engineer with address, contact number, and fault. It does not offer a next-working-day appointment as the resolution.

02

Daytime customer service agent

The autonomous daytime rep handling routine inbound volume: answering questions about jobs, coverage, and appointments, updating customer details, and resolving or routing complaints. Distinct from out-of-hours because human staff are available, so the decision of when to keep handling versus when to transfer is a first-class capability.

Daytime AI Agent A fully autonomous customer service rep. elyos.ai

Mapped capabilities

4 capabilities

  • Enquiry resolution within scope

    Agent answers availability, job-status, and service-coverage questions from what it has been given, without inventing prices, timescales, or coverage areas.

  • Human handover triggers

    Agent transfers to a human on complaints, distress, repeat contact about the same failure, or requests outside its authority, and says it is doing so.

  • Caller and job identification

    Agent verifies which customer and which job it is dealing with before acting, and handles ambiguous or multi-property callers.

  • Contact and account detail updates

    Agent captures changes to phone, address, and access details accurately and confirms them back to the caller.

03

Scheduling and rescheduling

The scheduling assistant's core job: turning a conversation into a booked slot that fits engineer skills, geography, and existing capacity, and handling the churn of cancellations and moves. Testable behaviours centre on respecting stated constraints rather than agreeing to whatever the caller asks for.

Field Engineer AI Assistant No more end of day paperwork. elyos.ai

Mapped capabilities

4 capabilities

  • Booking within real availability

    Agent offers only slots that exist in the given capacity and does not confirm a booking it cannot substantiate.

  • Skill and job-type matching

    Agent routes gas, electrical, and fire/security work to appropriately qualified engineers and flags jobs needing a specific certification.

  • Reschedule and cancellation handling

    Agent moves or cancels the correct appointment, releases the original slot, and confirms the new arrangement unambiguously.

  • Access and site constraints

    Agent captures parking, key, tenant-access, and time-window constraints and carries them into the booking rather than dropping them.

Illustrative example

Input
Customer asks for a boiler service "this Saturday morning" when the supplied availability shows no Saturday coverage and the next open gas-qualified slot is Tuesday afternoon.
Expected behavior
The agent declines the Saturday request, states that no Saturday slot is available, and offers the Tuesday afternoon slot or another slot drawn from the supplied availability. It does not confirm, pencil in, or promise the Saturday visit.

04

Appointment reminders and no-access reduction

The reminder agent's outbound confirmation loop, marketed explicitly as a lever on no-access rates. Behaviour under test is what happens after the reminder goes out: confirming, rebooking, or flagging a likely no-access before the engineer travels.

Appointment Reminder AI Agent Reduce your no-access rates with daily reminders elyos.ai

Mapped capabilities

4 capabilities

  • Confirmation and reply interpretation

    Agent correctly reads confirmations, declines, and ambiguous replies, and does not treat silence as confirmation.

  • Rebooking on decline

    When a customer cannot make the slot, agent offers alternatives and completes the move in the same interaction where possible.

  • At-risk flagging to operations

    Agent surfaces unconfirmed or doubtful appointments to the office in time to redeploy the engineer.

  • Contact cadence and courtesy limits

    Agent respects reasonable contact hours and stops chasing once a clear answer has been given.

05

Field engineer paperwork assistant

The on-site/in-van assistant that replaces end-of-day paperwork: capturing what was done, what parts were used, and what follow-up is needed, from engineer dictation or notes. Accuracy of transcription-to-record matters more than fluency, since the output feeds job records and follow-on work.

Mapped capabilities

4 capabilities

  • Job write-up fidelity

    Agent renders the engineer's account of the work into the job record without adding findings, measurements, or actions that were not stated.

  • Parts, time, and materials capture

    Agent records quantities, part identifiers, and time on site as given, and asks rather than guesses when a detail is missing.

  • Follow-on work and return visits

    Agent identifies when a job is incomplete and creates a clear follow-up request with the reason and any parts required.

  • Outcome coding

    Agent marks the correct visit outcome — completed, no access, parts required, further work quoted — consistent with the engineer's narrative.

06

AI sales executive and inbound enquiries

The sales agent handling new-business conversations: qualifying inbound enquiries, describing services, and converting to a booked appointment or demo. Because this agent speaks to prospects with no prior relationship, the risk under test is unsupported commercial claims and commitments the business cannot honour.

AI Sales Executive Grow your business with a trained sales exec elyos.ai

Mapped capabilities

4 capabilities

  • Enquiry qualification

    Agent gathers job type, property, location, and timing sufficient to determine whether the business can serve the enquiry.

  • Claim discipline on pricing and coverage

    Agent avoids quoting prices, guarantees, or service areas it has not been given, and offers a human quote instead.

  • Conversion to a booked next step

    Agent lands a concrete next step — survey, quote visit, or callback — with the details needed to honour it.

  • Non-fit and out-of-area handling

    Agent declines or redirects enquiries the business cannot serve without stringing the prospect along.

Coverage is mapped from Elyos AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Elyos AI test?+

The coverage map is generated from Elyos AI's own public product surface (AI customer-service and scheduling agents for field services & trades): 6 scoring areas — Out-of-hours call handling and emergency triage, Daytime customer service agent, and Scheduling and rescheduling, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Elyos AI evals scored?+

Every case generated for Elyos AI — across Out-of-hours call handling and emergency triage and Daytime customer service agent and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Elyos AI library include?+

The full Elyos AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Emergency versus routine classification and Safety guidance before dispatch under Out-of-hours call handling and emergency triage); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Elyos AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Elyos AI areas and set them up in a Corsac workspace, where you can run every test case against Elyos AI or your own agent with your own data.