All evals
C

Eval directory

Evals for Caretta

Eval coverage for Caretta, mapped from its public product surface.

About Caretta

Caretta is a real-time AI assistant for sales calls that connects a company's apps (CRMs, ERPs, LMSs, knowledge bases, docs) into a central "organisational intelligence" called Caretta Nous. During calls it takes notes and surfaces answers to account executives, and it also lives in the team's teamspace to push notes, answer questions, and run analyses and briefs. The company announced a $1.3m pre-seed round led by Y Combinator in December 2025.

Industry

realtime AI sales call assistant / sales organisational intelligence

Headquarters

Wilmington, Delaware, United States (stated place of business; open roles listed in Amsterdam and San Francisco)

Use the eval library for Caretta

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Caretta?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Realtime In-Call Assistance

Behavior while a sales call is in progress: surfacing the right answer to the rep at the right moment, handling technical and objection-heavy questions, and staying useful without derailing the conversation.

Caretta is native to your teamspace, and joins there 24/7 to push notes, answer questions, conduct analysis www.caretta.so

Mapped capabilities

4 capabilities

  • Answer surfacing on buyer questions

    Retrieves and presents a relevant, grounded answer when the buyer raises a product, technical, or pricing question during the call.

  • Objection handling support

    Supports the rep on hard objections about complex products, feature depth, and competitive pushback.

  • Live note capture

    Takes structured notes during the call covering what was asked, committed, and agreed.

  • Restraint and timing

    Avoids surfacing content when the call context does not call for it, so assistance does not overwhelm the rep.

Illustrative example

Input
Mid-call, a technical buyer asks whether the product syncs with their Odoo ERP. Connected docs describe an Odoo integration; the CRM account record does not mention it.
Expected behavior
Caretta surfaces the affirmative Odoo answer to the rep during the call, drawn from the connected docs, and attributes it to that source rather than to the CRM record.

02

Caretta Nous — Organisational Knowledge

The central brain assembled from connected apps (CRMs, ERPs, LMSs, knowledge bases, docs). Covers retrieval quality across sources, conflict resolution, and keeping the knowledge fresh as sources change.

Connect any app—CRMs, ERPs, LMSs, knowledge bases, docs and more—to create your team's very own central organisational intelligence. www.caretta.so

Mapped capabilities

4 capabilities

  • Cross-source retrieval

    Pulls the correct answer when the supporting fact lives in a doc, knowledge base, or CRM record rather than one obvious source.

  • Conflicting-source resolution

    Handles cases where two connected sources disagree, rather than silently picking one.

  • Freshness and staleness

    Reflects updated source content and flags knowledge that may be out of date.

  • Unknown-answer behavior

    States that the organisational knowledge does not cover a question instead of filling the gap.

Illustrative example

Input
During a live call, the buyer asks: "Are you SOC 2 Type II certified, and can you send the report?" No connected source contains any certification record.
Expected behavior
Caretta tells the rep the organisational knowledge contains no certification record, rather than asserting or denying certification. It suggests routing the question to a follow-up instead of supplying a number, date, or document.

03

Teamspace Agent Workflows

Caretta's always-on presence in the team's teamspace: pushing notes after calls, answering questions asynchronously, running analyses, and producing morning briefs.

As Caretta takes notes during call, it also reveals information and answers to reps when needed www.caretta.so

Mapped capabilities

4 capabilities

  • Post-call note push

    Delivers call notes into the teamspace in a form the team can act on.

  • Async question answering

    Answers teamspace questions using the same organisational knowledge available in-call.

  • Morning briefs

    Assembles a scheduled brief from account and call activity.

  • Requested analyses

    Runs the analysis a team member asks for and states what data it drew on.

04

App Connections & Integration Coverage

Connecting and working with the customer's stack — Zoom, Google Meet, Salesforce, HubSpot, Notion, Google Drive, Slack, Microsoft Teams, Cal.com, Telegram, Odoo, OneDrive — including behavior when a connection is missing or degraded.

Mapped capabilities

4 capabilities

  • Meeting platform coverage

    Operates across the supported call platforms without behavior differing in ways the rep must compensate for.

  • CRM read and write-back

    Uses CRM account context and pushes call output back in structured form.

  • Missing or unconnected source

    Behaves predictably when a relevant app is not connected and says so rather than guessing.

  • Connector failure handling

    Degrades gracefully when a connected app is unavailable mid-workflow.

05

Grounding, Accuracy & Failure Recovery

The reliability discipline the company frames as part of the product: grounded answers, no fabricated facts on live calls, and clean recovery when something breaks mid-call.

Mapped capabilities

4 capabilities

  • No fabricated product claims

    Declines to assert capabilities, numbers, or commitments not present in connected sources.

  • Attribution to source

    Ties a surfaced answer back to where the knowledge came from.

  • Mid-call degradation

    Recovers or fails visibly if realtime assistance drops during a call, rather than silently going quiet.

  • Note fidelity

    Notes reflect what was actually said, without inventing commitments.

06

Account, Trial & Terms Surface

Signup and commercial surface described publicly: Google-based signup, a seven-day free trial with cancel-anytime, an administrative account user, and the subscription/billing terms.

7 days free · Cancel anytime www.caretta.so

Mapped capabilities

4 capabilities

  • Signup and onboarding

    Google sign-in and initial setup path for a new team.

  • Trial and cancellation terms

    Represents the seven-day free trial and cancel-anytime terms accurately when asked.

  • Admin account role

    Administrative user setup and account-level control described in the terms.

  • Use restrictions

    Handles requests that conflict with the stated use restrictions in the terms.

Coverage is mapped from Caretta's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Caretta test?+

The coverage map is generated from Caretta's own public product surface (realtime AI sales call assistant / sales organisational intelligence): 6 scoring areas — Realtime In-Call Assistance, Caretta Nous — Organisational Knowledge, and Teamspace Agent Workflows, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Caretta evals scored?+

Every case generated for Caretta — across Realtime In-Call Assistance and Caretta Nous — Organisational Knowledge and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Caretta library include?+

The full Caretta library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Answer surfacing on buyer questions and Objection handling support under Realtime In-Call Assistance); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Caretta or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Caretta areas and set them up in a Corsac workspace, where you can run every test case against Caretta or your own agent with your own data.