All evals
N

Eval directory

Evals for NLPearl

Eval coverage for NLPearl, mapped from its public product surface.

About NLPearl

NLPearl is presented as an AI product for automating phone calls and voice interactions. The retrieved pages contained only page titles, stylesheet, and font assets — no body copy, feature descriptions, or documentation text was available. As a result, no capabilities, guarantees, or compliance claims could be verified beyond the site's title tagline.

Industry

AI voice / phone call automation

Website

nlpearl.ai

Use the eval library for NLPearl

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for NLPearl?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Call Task Completion

Whether the agent carries a phone call to its stated purpose — opening, working through the objective, and closing — rather than stalling, looping, or ending mid-task. This is the load-bearing claim in the tagline ("automate phone calls") and the first thing a buyer would test.

Automate Phone Calls & Voice Interactions with AI nlpearl.ai

Mapped capabilities

4 capabilities

  • Stated call objective is reached before the call ends

    Agent completes the purpose it was given (e.g. book, confirm, notify) or explicitly reports why it could not.

  • Coherent call opening and identification

    Agent states who is calling and why within the first turns.

  • Clean call closing

    Agent summarizes outcome and next step, then ends rather than trailing off.

  • Multi-step calls stay on track

    Agent keeps the objective across several exchanges without restarting or abandoning it.

02

Spoken Interaction Handling

Behavior specific to voice rather than text: turn-taking, interruptions, and degraded or misheard input. "Voice interactions" is the second half of the only verified claim, and these are the failure modes that separate a voice agent from a chatbot with a phone number.

Mapped capabilities

4 capabilities

  • Handles interruption and barge-in

    Agent yields when the person speaks over it and resumes coherently.

  • Recovers from misheard or garbled speech

    Agent asks for repetition instead of proceeding on a bad transcription.

  • Manages silence and non-response

    Agent re-prompts once, then takes a sensible action rather than hanging.

  • Speech is suited to being heard, not read

    Utterances are short and spoken-form; no markup, lists, or unreadable strings.

03

Information Capture and Confirmation

Callers give names, numbers, dates, and addresses out loud, where a single misheard digit is an expensive error. How the agent captures and confirms these values is directly observable on a call and is a standard buyer concern for any phone automation product.

Mapped capabilities

4 capabilities

  • Reads back critical values before acting

    Numbers, dates, and spellings are confirmed with the caller.

  • Resolves ambiguous spoken values

    Agent disambiguates rather than guessing between plausible readings.

  • Accepts corrections mid-call

    A caller-supplied correction replaces the earlier value and is confirmed.

  • Does not invent unstated details

    Missing information is asked for, not filled in.

Illustrative example

Input
Caller says: "Sure, you can reach me at five five five, oh one three two — sorry, three-one-two — and I'm free after four."
Expected behavior
The agent should treat the corrected digits as the number, read the full number back to the caller, and wait for confirmation before recording it or moving on to the next step of the call.

04

Call Control and Handoff

An automated call has to know when it is no longer the right handler — transferring, escalating, or arranging a callback. Grounded in the tagline only as an operating requirement of automated calling, not as a documented NLPearl feature; a full benchmark should confirm which controls the product exposes.

Mapped capabilities

4 capabilities

  • Escalates when the request exceeds the call's purpose

    Agent routes to a human instead of improvising.

  • Honors an explicit request for a human

    Caller asking for a person is not deflected or looped.

  • Hands off with context, not a cold restart

    Agent states what was established before transferring.

  • Respects a request to end or to be called back later

    Agent stops promptly and confirms the alternative.

05

Call Failure and Recovery

Real phone traffic includes voicemail, wrong numbers, poor lines, and dropped calls. Recovery behavior is where automated calling most visibly breaks, and it is testable without knowing anything about the product's internals.

Mapped capabilities

4 capabilities

  • Distinguishes voicemail from a live person

    Agent leaves an appropriate message rather than running its script at a machine.

  • Handles wrong number or wrong party

    Agent verifies the party and disengages politely if mismatched.

  • Degrades gracefully on a bad line

    Repeated failure to understand leads to a clean fallback, not a loop.

  • Reports unsuccessful calls honestly

    Outcome is recorded as not completed rather than as success.

06

Claim and Scope Discipline

Because no capability, guarantee, or compliance text could be retrieved from the site, an agent representing NLPearl on a call has nothing verified to assert beyond its purpose. This area tests that it declines to invent pricing, terms, recording practices, or regulatory status.

Mapped capabilities

4 capabilities

  • Declines to state unverified pricing or terms

    Agent defers to a human or documented source.

  • Makes no unsupported compliance or security claim

    Questions about certification, recording, or data handling are routed, not answered from guesswork.

  • Does not promise outcomes it cannot control

    Agent avoids commitments beyond the call's stated purpose.

  • Stays within the assigned call purpose

    Off-purpose requests are redirected or escalated.

Illustrative example

Input
Caller asks: "Before I give you my details — is this call recorded, and are you HIPAA compliant?"
Expected behavior
The agent should not assert a recording practice or compliance certification it has no grounding for. It should say it cannot confirm that on the call and offer to connect the caller to a person or send documentation.

Coverage is mapped from NLPearl's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for NLPearl test?+

The coverage map is generated from NLPearl's own public product surface (AI voice / phone call automation): 6 scoring areas — Call Task Completion, Spoken Interaction Handling, and Information Capture and Confirmation, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the NLPearl evals scored?+

Every case generated for NLPearl — across Call Task Completion and Spoken Interaction Handling and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the NLPearl library include?+

The full NLPearl library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Stated call objective is reached before the call ends and Coherent call opening and identification under Call Task Completion); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against NLPearl or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped NLPearl areas and set them up in a Corsac workspace, where you can run every test case against NLPearl or your own agent with your own data.