All evals
Bland AI

Eval directory

Evals for Bland AI

Eval coverage for Bland AI, mapped from its public product surface.

About Bland AI

Bland is an enterprise voice AI platform for building and running AI phone agents, aimed at regulated industries such as healthcare, insurance, financial services, and logistics. It builds its own in-house stack (voice, LLM, TTS, STT) and runs self-hosted so customer data does not pass through third parties. Agents can be built from a prompt via its Norm builder, tested against scenarios before launch, and deployed across voice, SMS, iMessage, and web chat with shared memory.

Employees

~100

Industry

enterprise voice AI platform for phone agents

Headquarters

San Francisco

Use the eval library for Bland AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Bland AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agent Building with Norm

Turning a plain-language description of a business need into a production-shaped agent draft, including pathway, voice, and connected tools.

Mapped capabilities

4 capabilities

  • Prompt-to-agent draft

    Produces a named agent draft with a stated use case from a one-line description.

  • Pathway construction

    Assembles the conversational pathway steps implied by the request.

  • Voice and persona selection

    Assigns a voice and agent persona appropriate to the described use case.

  • Integration wiring

    Attaches the calendar, CRM, SMS, and transfer components the request implies.

02

Pre-Launch Scenario Testing

Running an agent against scenarios and edge cases before production, and reporting what passes, fails, and regresses.

Mapped capabilities

4 capabilities

  • Edge-case scenario coverage

    Generates and runs scenarios beyond the happy path for a given agent.

  • Pass/fail reporting

    Reports failing scenarios distinctly from passing ones with the failure cause.

  • Reruns after fixes

    Re-executes previously failing scenarios and reports fixed versus still-failing.

  • Staging-to-production gating

    Keeps an agent in staging rollout until tests are clean.

03

Live Call Control and Actions

Dispatching and handling real phone calls, including in-call external actions, escalation, and high-volume campaigns.

“We have the lowest latency on the planet, making your calls sound human.” docs.bland.ai

Mapped capabilities

4 capabilities

  • Outbound and inbound dispatch

    Sends outbound calls and configures inbound numbers for support use.

  • Live API calls during a call

    Invokes an external API mid-call and uses the result in the conversation.

  • Warm transfer to a human

    Escalates to a human with context rather than dropping the caller.

  • Batch call campaigns

    Launches large batches of calls within configured limits.

04

Omnichannel Continuity

One agent operating across voice, SMS, iMessage, and web chat while carrying the same customer memory between channels.

Mapped capabilities

4 capabilities

  • Shared memory across channels

    Recalls prior voice-call facts when the same customer messages later.

  • Channel handoff

    Continues an in-flight workflow when the customer switches channel.

  • Document and artifact delivery

    Sends the correct document into the thread at the right workflow milestone.

  • Embedded web and chat agents

    Runs the same agent embedded in a customer-facing web application.

Illustrative example

Input
Caller locks a 6.125% refinance rate on a voice call. The next day the same customer texts: "What are my float-down options?"
Expected behavior
The SMS reply reflects the existing lock at 6.125% from the prior voice call and answers float-down in that context, without asking the customer to restate their rate, loan, or identity.

05

Security, Compliance, and Deployment Controls

The controls regulated buyers evaluate: self-hosted data handling, deployment topology, contractual guarantees, and resistance to prompt manipulation.

“Bland is fully self hosted, meaning your data stays entirely secure from end-to-end.” docs.bland.ai

Mapped capabilities

4 capabilities

  • Self-hosted data handling

    Explains that data does not pass through third-party model providers.

  • On-prem and VPC deployment

    Describes the deployment options available on the Enterprise plan.

  • Enterprise compliance controls

    Covers BAA, SSO, and data residency availability accurately.

  • Version lock and model stability

    Explains version lock and that models do not change underneath a customer.

06

Plans, Limits, and Billing Explanation

Answering questions about rates, caps, and included capacity correctly, and routing a customer to the right plan.

Mapped capabilities

4 capabilities

  • Per-minute and transfer rates

    Quotes the correct talk-time and transfer rate for a named plan.

  • Concurrency and volume caps

    States the correct concurrent, hourly, and daily call limits per plan.

  • Knowledge base and voice limits

    States the correct knowledge base and voice counts per plan.

  • Billing model explanation

    Explains that LLM, STT, and TTS are included with no token charges.

Illustrative example

Input
We're on the Build plan and want to run a campaign at 80 concurrent calls tomorrow. Can we do that on our current plan?
Expected behavior
States that Build allows 50 concurrent calls and 2,000 calls per day, so 80 concurrent exceeds the plan, and points to Scale (100 concurrent) or Enterprise as the path to that volume.

Coverage is mapped from Bland AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Bland AI test?+

The coverage map is generated from Bland AI's own public product surface (enterprise voice AI platform for phone agents): 6 scoring areas — Agent Building with Norm, Pre-Launch Scenario Testing, and Live Call Control and Actions, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Bland AI evals scored?+

Every case generated for Bland AI — across Agent Building with Norm and Pre-Launch Scenario Testing and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Bland AI library include?+

The full Bland AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Prompt-to-agent draft and Pathway construction under Agent Building with Norm); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Bland AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Bland AI areas and set them up in a Corsac workspace, where you can run every test case against Bland AI or your own agent with your own data.