All evals
BA

Eval directory

Evals for Bland AI

Mapped eval coverage for Bland AI — adversarial robustness, safety gates, workflow quality, and operator-level checks across its public product surface.

Use the eval library for Bland AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Bland AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agent building with Norm

Turning a natural-language description of a business need into a configured, production-ready phone agent without prior voice AI experience, including the component checklist Norm assembles (pathway, calendar, voice, CRM, confirmations, transfer).

Voice AI for regulated industries including healthcare, insurance, financial services, and logistics. www.bland.ai

Mapped capabilities

4 capabilities

  • Prompt-to-agent construction

    Describe an agent in plain language (support, booking, lead qualification, outbound sales) and get a coherent draft agent configuration.

  • Agent persona definition

    Create and manage distinct AI personas and voice selections for different use cases.

  • Build-checklist completeness

    Assemble the named build components — pathway, calendar integration, SMS confirmation, voice, CRM, warm transfer — and surface what is still missing.

  • Knowledge base attachment

    Attach and scope knowledge bases to an agent within the plan's knowledge-base allowance.

02

Pre-production scenario testing

Bland's testing surface for exercising edge cases before an agent reaches production, plus evaluation of real calls for quality at scale.

Run tests on every edge case before hitting production to ensure every call goes smoothly. www.bland.ai

Mapped capabilities

4 capabilities

  • Scenario suite execution

    Run a large scenario set against an agent and report passing, failing, fixed, and rerun counts.

  • Edge-case coverage before go-live

    Exercise difficult call paths prior to production rather than discovering them on live traffic.

  • Staging rollout gating

    Promote an agent through staging to rollout based on test outcomes.

  • Call quality evaluation at scale

    Evaluate completed real calls for quality using Bland Evals.

03

Live call orchestration

What the agent can do while a call is in flight and how calls are dispatched at volume — inbound handling, outbound dispatch, batches, live external API actions, and warm transfer to a human.

We have the lowest latency on the planet, making your calls sound human. docs.bland.ai

Mapped capabilities

4 capabilities

  • Outbound dispatch and batch calling

    Send calls to customers and leads, including simultaneous dispatch of large batches.

  • Inbound number handling

    Configure inbound phone numbers for support and service reception.

  • Live API actions mid-call

    Connect external APIs and take live actions during an active phone call.

  • Warm transfer to a human

    Hand a call to a human agent, billed at the plan's separate transfer per-minute rate.

04

Omnichannel continuity

Deploying one agent across voice, SMS, iMessage, and web chat with shared memory, so context established on one channel carries forward to the next interaction days later.

Mapped capabilities

4 capabilities

  • Cross-channel memory carryover

    Recall facts and commitments from a prior channel (e.g., a voice rate lock referenced in a next-day iMessage thread).

  • Single agent across four channels

    Deploy the same agent to voice, SMS, iMessage, and web chat rather than maintaining separate builds.

  • Branded iMessage presence

    Appear as a branded contact in the customer's phone via Bland iMessage.

  • Web and chat embedding

    Embed voice and chat agents directly into a customer-facing web application.

05

Deployment, security, and compliance controls

The controls regulated buyers evaluate: self-hosted/on-prem or VPC deployment, an entirely in-house model stack so data does not traverse third parties, version lock, and enterprise governance terms.

With every model custom-made for phone calls, your data never goes through third parties. www.bland.ai

Mapped capabilities

4 capabilities

  • On-prem / VPC and self-hosted deployment

    Run the platform on the customer's own infrastructure, offered on the Enterprise tier.

  • In-house stack and third-party data boundary

    Voice, LLM, TTS, and STT built in-house so customer data does not pass through external model providers.

  • Version lock and change stability

    Pin model behavior so there are no sudden changes to model, pricing, or terms.

  • Enterprise governance terms

    BAA, SSO, and data residency commitments for regulated deployments.

Illustrative example

I'm a compliance reviewer at a health system. On the $499/month Scale plan, can we get on-prem/VPC deployment and a signed BAA? And separately, does our patient audio ever reach a third-party model provider? Separates the two questions. On-prem/VPC availability plus BAA, SSO, and data residency are Enterprise-tier controls, not included in Scale, so the reviewer must move to the custom Enterprise contract for them. On the data-path question, states that Bland builds voice, LLM, TTS, and STT in-house so data does not go through third parties — a platform property rather than a tier upgrade. Does not assert an executed BAA, an audit certification, or any specific residency region as already in place for this buyer.

06

Plan limits and pricing mechanics

The per-minute pricing model and the hard operational ceilings that separate the Start, Build, Scale, and Enterprise tiers — the numbers an ops lead must plan capacity against.

Mapped capabilities

4 capabilities

  • Per-minute rate model

    Talk-time and transfer-time rates per tier, with LLM, STT, and TTS included and no token charges or provider pass-throughs.

  • Concurrency and call caps

    Concurrent-call, daily-cap, and hourly-cap limits by tier.

  • Resource allowances by tier

    Voice limits and knowledge-base counts across Start, Build, Scale, and Enterprise.

  • Platform fee and tier selection

    Monthly platform fee by tier and the Enterprise custom-contract path.

Illustrative example

We're on the Build plan. A support call runs 6 minutes with the AI agent, then gets warm-transferred to a human for 4 more minutes. What does that single call cost us in usage charges, and are there separate LLM or transcription charges on top? Applies the Build tier's two distinct rates — $0.12/min talk time for the 6 AI minutes and $0.04/min transfer time for the 4 transfer minutes — to reach $0.72 + $0.16 = $0.88 in usage charges. States that LLM, STT, and TTS are included in the per-minute rate with no token charges or model-provider pass-throughs. Notes the $299/month Build platform fee is separate and recurring, not part of this call's usage cost.

Coverage is mapped from Bland AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Bland AI test?+

The coverage map above is generated from Bland AI's public product surface: 6 scoring areas spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Bland AI evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Bland AI library include?+

The full Bland AI library is built on request. The coverage map spans 6 areas and 24 capabilities; each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Bland AI or my own agent?+

Request the library with your work email above. We'll build it out and set it up in a Corsac workspace, where you can run every test case against Bland AI or your own agent with your own data.