All evals
T

Eval directory

Evals for Thoughtly

Eval coverage for Thoughtly, mapped from its public product surface.

About Thoughtly

Thoughtly is an AI engagement platform that gives CRMs a voice: agents call, text, and email inbound leads automatically and follow up across channels until the lead responds. It emphasizes sub-60-second speed-to-lead, branded caller ID, dynamic qualification, and warm handoff to human reps. It sells in three tiers (Flex, Scale, Enterprise) with two-way CRM sync, 200+ integrations, and compliance features that scale by tier.

Industry

omnichannel AI voice/SMS/email agent platform for GTM and revenue teams

Use the eval library for Thoughtly

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Thoughtly?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Speed-to-Lead Intake & Triggering

Whether an inbound signal turns into a live outreach attempt fast enough and from the right entry point. Covers the sub-60-second promise, the variety of supported triggers, and correct handling of leads that arrive outside a clean form fill.

Mapped capabilities

4 capabilities

  • Sub-60-second first contact on form fill

    Form submission triggers a voice attempt within the stated speed-to-lead window rather than queuing.

  • Trigger source coverage

    Form fill, webhook, list membership, pipeline stage, and workflow events each initiate a run.

  • Channel selection at first touch

    Choosing voice versus text for the opening attempt based on the lead's intent and contact data.

  • Uncontacted and returning-visitor backlog

    Working the bottom of the list — returning visitors and never-contacted contacts — not just top leads.

02

Voice Agent Conversation & Qualification

The quality of the live call itself: sounding human, staying on the trained script, adapting questions to answers, and ending in a correct disposition. This is the surface Thoughtly leads with and where conversion is won or lost.

CRM-driven AI voice agents that call your leads — then text, follow up thoughtly.com

Mapped capabilities

4 capabilities

  • Dynamic qualification that adapts to answers

    Follow-up questions branch on what the lead said instead of replaying a fixed script.

  • Warm transfer with full context

    Handoff fires at the right moment and carries captured fields and call context to the human rep.

  • In-call booking against live availability

    Offering and confirming real scheduler slots without double-booking or inventing times.

  • Branded caller ID and identity claims

    Agent identifies itself and the brand consistently with the verified caller identity presented.

Illustrative example

Input
On a live inbound call, the lead states a $620,000 loan amount, a 30-year fixed preference, and asks to speak with a loan officer now.
Expected behavior
The agent treats the lead as qualified, tells them it is connecting a loan officer, and transfers. The transfer payload carries the loan amount, product type, and contact details so the rep opens on context rather than re-asking qualification questions.

03

Cross-Channel Follow-Up Orchestration

One agent persisting across voice, SMS/iMessage/WhatsApp, and email until the lead responds. Covers cadence judgment, context continuity between channels, and recovery when a channel fails.

Mapped capabilities

4 capabilities

  • Context continuity across channels

    The text and email reference the prior call thread rather than restarting the conversation.

  • Voice-to-SMS failover on missed calls

    An unanswered or dropped call pivots to messaging instead of dead-ending.

  • Cadence and back-off judgment

    Knowing when to push again and when to stop, per the intent-driven follow-up model.

  • Per-lead email drafting and threaded replies

    Personalized email that stays in-thread instead of templated first-name merges.

04

CRM & Integration Sync Fidelity

Whether the system of record stays true after the agent acts. Covers two-way sync with the named CRMs, write-back of conversation artifacts, scheduler integration, and auth boundaries.

OAuth on every connection. No credentials shared with the agent. thoughtly.com

Mapped capabilities

4 capabilities

  • Write-back of summary, disposition, and transcript

    Call artifacts land on the correct contact, lead, or deal record.

  • Two-way state reconciliation

    Reading pipeline state and writing outcomes without clobbering CRM-side edits.

  • Scheduler round-robin and routing rules

    Calendly, Cal.com, and Acuity routing rules carry through into the booking the agent makes.

  • OAuth boundaries on connections

    Native auth is used and raw credentials are never surfaced to or requested by the agent.

06

Plan Tiers, Entitlements & Support Boundaries

Thoughtly gates volume, compliance artifacts, and support by tier across Flex, Scale, and Enterprise. Correctly representing and enforcing those boundaries matters to both buyers and deployed agents.

Mapped capabilities

4 capabilities

  • Tier-gated capability claims

    Custom LLM, HIPAA/BAA, SSO/SCIM, and data residency are attributed to Enterprise, not lower tiers.

  • Concurrency and volume ceilings

    Behavior at and beyond the concurrent-call limits associated with the customer's plan.

  • Compliance artifact availability by tier

    SOC 2 Type II reports and TCPA review offered only where the plan actually includes them.

  • Support and escalation path accuracy

    Routing to the support channel the plan entitles — email, Slack/CSM, or named TAM hotline.

Coverage is mapped from Thoughtly's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Thoughtly test?+

The coverage map is generated from Thoughtly's own public product surface (omnichannel AI voice/SMS/email agent platform for GTM and revenue teams): 6 scoring areas — Speed-to-Lead Intake & Triggering, Voice Agent Conversation & Qualification, and Cross-Channel Follow-Up Orchestration, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Thoughtly evals scored?+

Every case generated for Thoughtly — across Speed-to-Lead Intake & Triggering and Voice Agent Conversation & Qualification and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Thoughtly library include?+

The full Thoughtly library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Sub-60-second first contact on form fill and Trigger source coverage under Speed-to-Lead Intake & Triggering); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Thoughtly or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Thoughtly areas and set them up in a Corsac workspace, where you can run every test case against Thoughtly or your own agent with your own data.