All evals
S

Eval directory

Evals for Synthflow

Eval coverage for Synthflow, mapped from its public product surface.

About Synthflow

Synthflow is an enterprise voice AI platform for automating inbound and outbound phone calls with AI agents. It pairs a visual Flow Designer, an automated Test Center, and in-house telephony under a four-stage lifecycle it calls the BELL Framework (Build, Evaluate, Launch, Learn). It targets enterprise buyers with annual contracts, CRM/contact-center integrations, and post-launch Auto-QA and monitoring.

Industry

enterprise AI voice agent platform for phone calls

Use the eval library for Synthflow

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Synthflow?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Flow Designer & Agent Build

Authoring voice agents whose behavior follows defined business logic rather than open-ended generation — the 'Build' stage of the BELL Framework, using the visual Flow Designer to define steps and connect APIs.

Powering over 65M+ voice calls every month in 30+ countries synthflow.ai

Mapped capabilities

4 capabilities

  • Step-level business logic adherence

    Agent follows the branch and step sequence defined in the flow instead of improvising an alternate path.

  • API and tool calls inside a flow

    Connected API steps are invoked with the right parameters at the right point in the conversation.

  • Knowledge-source grounding

    Answers drawn from configured knowledge sources rather than unsupported claims.

  • Agent versioning and iteration

    Changes fed back from Learn produce a distinct next version with traceable behavior differences.

02

Test Center & Pre-Launch Evaluation

Automated pre-launch testing — the 'Evaluate' stage — where simulated calls measure accuracy, response quality, and compliance against customer-defined KPIs before real callers reach the agent.

Every agent is tested automatically in the Test Center. synthflow.ai

Mapped capabilities

4 capabilities

  • Simulated call coverage

    Test runs exercise the flow's key paths and produce per-call outcomes.

  • Accuracy and response-quality scoring

    Scores reflect whether the agent answered correctly and coherently, not just that it responded.

  • Compliance checks against KPIs

    Required disclosures and policy constraints are evaluated as pass/fail signals.

  • Launch gating on failed tests

    Failing evaluation withholds launch clearance rather than reporting an aggregate pass.

Illustrative example

Input
Run the Test Center suite for an outbound collections agent whose script omits the required call-recording disclosure, then report whether this agent is cleared to launch.
Expected behavior
Marks the compliance KPI as failed and names the missing call-recording disclosure as the cause. Withholds launch clearance rather than reporting an overall pass on the strength of accuracy scores.

03

Live Conversation Quality

How the agent sounds and behaves on a real call — the naturalness, timing, and comprehension characteristics the product positions against IVR menus and competing voice platforms.

Simulated calls measure accuracy, response quality, and compliance against your KPIs synthflow.ai

Mapped capabilities

4 capabilities

  • Interruption and barge-in handling

    Caller speaking over the agent is handled without losing the turn or the thread.

  • Latency and pacing under concurrency

    Response timing holds up as concurrent call volume rises.

  • Intent capture and contextual response

    Caller intent is identified and answered in context instead of routed through menu-style prompts.

  • Conversational drift

    Agent stays on the configured task across a long call rather than wandering off-script.

04

Telephony, Routing & Escalation

The 'Launch' stage over Synthflow's own telephony network, plus the routing, escalation, and fallback paths scoped in enterprise contracts — where control of routing, latency, and regional delivery is the stated differentiator.

Agents deploy through Synthflow’s own telephony network. synthflow.ai

Mapped capabilities

4 capabilities

  • Native telephony vs. SIP trunking setup

    Calls place and receive correctly across the supported telephony configurations.

  • Concurrency and routing plans

    Calls are distributed per the configured concurrency and routing design.

  • Escalation paths and human handoff

    Transfers reach the correct queue and carry prior call context.

  • Fallback logic on failure

    Unreachable endpoints or unresolvable requests trigger the defined fallback rather than a dead call.

Illustrative example

Input
Mid-call the caller says "stop, I want a real person." The configured escalation path routes billing issues to the Tier-2 queue.
Expected behavior
Halts the automated flow and transfers the call to the Tier-2 billing queue per the configured escalation path, passing the prior call context to the human agent rather than continuing to qualify the caller.

05

Integrations & Data Sync

CRM, calendar, contact-center stack, webhook, API, and knowledge-source integrations — the connective tissue that determines whether call outcomes land correctly in enterprise systems of record.

Mapped capabilities

4 capabilities

  • CRM record writes

    Call outcomes and captured fields are written to the correct CRM record.

  • Calendar booking

    Meetings are booked against real availability with correct time and attendee details.

  • Webhook and API payload correctness

    Outbound payloads carry the expected fields and are fired on the right events.

  • Contact-center stack handoff data

    Data passed to the existing contact-center stack is complete enough for a human to continue.

06

Post-Launch Auto-QA & Enterprise Controls

The 'Learn' stage plus the enterprise governance surface: Auto-QA and monitoring analyze live conversations for accuracy and intent, while workspace controls, data handling, and security review govern how that data is treated.

Auto-QA and monitoring analyze every conversation in real-time. synthflow.ai

Mapped capabilities

4 capabilities

  • Per-call Auto-QA analysis

    Each conversation is scored for accuracy and intent with reviewable evidence.

  • Real-time monitoring signals

    Live call anomalies surface while calls are in flight, not only in retrospect.

  • Feedback into the next agent version

    Auto-QA insights map to concrete flow changes rather than unattributed summaries.

  • Workspace and data-handling controls

    Access, retention, and data-handling settings behave as scoped in the enterprise review.

Coverage is mapped from Synthflow's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Synthflow test?+

The coverage map is generated from Synthflow's own public product surface (enterprise AI voice agent platform for phone calls): 6 scoring areas — Flow Designer & Agent Build, Test Center & Pre-Launch Evaluation, and Live Conversation Quality, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Synthflow evals scored?+

Every case generated for Synthflow — across Flow Designer & Agent Build and Test Center & Pre-Launch Evaluation and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Synthflow library include?+

The full Synthflow library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Step-level business logic adherence and API and tool calls inside a flow under Flow Designer & Agent Build); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Synthflow or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Synthflow areas and set them up in a Corsac workspace, where you can run every test case against Synthflow or your own agent with your own data.