All evals
P

Eval directory

Evals for Phonely

Eval coverage for Phonely, mapped from its public product surface.

About Phonely

Phonely is a voice AI platform for building and deploying AI phone agents that answer, route, and resolve customer calls. Teams pick a voice, train the agent on a knowledge base or website, connect CRM and scheduling tools, then deploy across phone, chat, SMS, and API with call recording, transcription, and AI analytics. Plans run from a free tier with 100 minutes to enterprise pricing with agent buildout, SIP trunking, and a HIPAA BAA.

Industry

voice AI phone answering / conversational AI platform

Use the eval library for Phonely

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Phonely?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agent Builder & Voice Setup

Configuring an agent before it takes live calls: picking a voice and language, cloning a voice, and training the agent on business knowledge.

Full Agent Buildout Fine tuning & SIP trunking Slack Channel Support HIPAA BAA www.phonely.ai

Mapped capabilities

4 capabilities

  • Voice, accent, and language selection

    Choosing among the advertised 100+/1,000+ voices and matching the language or accent the buyer asks for.

  • Voice cloning setup

    Guiding a user through cloning their own voice from the platform, including consent and quality prerequisites.

  • Knowledge base and website training

    Ingesting an uploaded knowledge base or connected website so the agent answers from business-specific content.

  • Agent buildout and deployment config

    Taking a configured agent from builder to a live number or channel, including enterprise full-buildout paths.

02

Live Call Handling & Conversation Quality

How the agent behaves in an active conversation: turn-taking, grounding in trained content, and filtering unwanted calls.

Mapped capabilities

4 capabilities

  • Turn-taking and interruption handling

    Natural handoff of speaking turns, recovery from caller interruptions and overlapping speech.

  • Knowledge-grounded answers

    Answering from the trained knowledge base and declining to invent facts it was not trained on.

  • Multilingual and accented conversation

    Holding the conversation in the caller's language and handling accent variation without degrading comprehension.

  • Spam and unwanted call filtering

    Identifying and filtering spam calls per plan-level spam filtering capability.

03

Routing, Transfer & Escalation

Getting the caller to the right destination: intent capture, warm versus cold handoff, and escalation to a human with context.

Mapped capabilities

4 capabilities

  • Intent capture before connect

    Reading caller intent in natural language and selecting the right department or queue.

  • Warm transfer with context

    Passing caller context and whisper summaries to the receiving agent so the caller does not repeat themselves.

  • Cold/blind transfer for simple routing

    Fast routing when no context handoff is required, and choosing it appropriately over warm transfer.

  • Human escalation triggers

    Recognizing complex, sensitive, or VIP calls that should leave AI containment.

Illustrative example

Input
Caller: "I've called twice about a double charge on my March invoice and nobody fixed it. Just put me through to a billing supervisor."
Expected behavior
The agent routes to billing escalation and passes the collected context — repeat contact, double charge, March invoice — to the receiving party before or as the call connects, rather than transferring blind and forcing the caller to restate the issue.

04

Integrations & Real-Time Task Execution

Actions the agent performs mid-call against connected systems, including tools that have no API.

Phonely integrates with your tools for real-time appointment booking, CRM updates, and so much more. www.phonely.ai

Mapped capabilities

4 capabilities

  • CRM lookup and update

    Pulling customer context in real time and writing back call outcomes without human intervention.

  • Appointment booking and rescheduling

    Checking availability in connected scheduling software and confirming or moving a booking.

  • Browser-based automation for API-less tools

    Driving legacy systems through prebuilt browser automations when no API exists.

  • API and custom integration behavior

    Advanced API integrations available on higher plans, including failure handling when a call to an external tool does not return.

Illustrative example

Input
Caller: "Book me for Thursday at 2pm." The connected scheduling tool returns no availability at that time and the write attempt does not succeed.
Expected behavior
The agent does not tell the caller the appointment is confirmed. It states the requested time is unavailable, offers an alternative slot from the calendar or a callback, and leaves no booking recorded in the scheduling system.

05

Omnichannel Delivery

Consistent agent behavior across the phone, chat, SMS, and API surfaces the platform exposes.

build and optimize AI agents across voice, chat, SMS, and more, from one conversation to one million www.phonely.ai

Mapped capabilities

4 capabilities

  • Voice and chat parity

    Same answers and policies whether the customer arrives by phone or chat.

  • SMS conversation and notifications

    Conversational SMS plus email/SMS notification behavior configured on the account.

  • API-invoked agent sessions

    Programmatic deployment of the agent as an API surface.

  • Cross-channel context continuity

    Carrying conversation context when a customer moves between channels.

06

Post-Call Analytics, Compliance & Plan Limits

What the platform produces after the call, and the policy boundaries around recordings, regulated data, and plan entitlements.

Mapped capabilities

4 capabilities

  • Recording, transcription, and summaries

    Producing accurate transcripts and AI summaries of completed calls.

  • Sentiment analysis and custom reports

    AI-driven insights and custom reports built in-platform or exported to a monitoring tool.

  • A/B testing of agent variants

    Comparing agent configurations against call outcomes.

  • Regulated data and plan entitlement boundaries

    Handling of PCI/HIPAA-sensitive caller data and accurate statements about minute allowances and tier-gated features.

Coverage is mapped from Phonely's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Phonely test?+

The coverage map is generated from Phonely's own public product surface (voice AI phone answering / conversational AI platform): 6 scoring areas — Agent Builder & Voice Setup, Live Call Handling & Conversation Quality, and Routing, Transfer & Escalation, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Phonely evals scored?+

Every case generated for Phonely — across Agent Builder & Voice Setup and Live Call Handling & Conversation Quality and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Phonely library include?+

The full Phonely library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Voice, accent, and language selection and Voice cloning setup under Agent Builder & Voice Setup); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Phonely or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Phonely areas and set them up in a Corsac workspace, where you can run every test case against Phonely or your own agent with your own data.