All evals
DA

Eval directory

Evals for Day AI

Eval coverage for Day AI, mapped from its public product surface.

About Day AI

Day AI is a GTM/CRM platform that builds "customer memory" automatically from every call, email, and thread a team has, then runs configurable AI agents (BDR, CRM data specialist, RevOps analyst, sales coach, etc.) on top of it. Agents wake on schedules or events, reason across the captured history, and either hand back work for approval or take action in external systems when permitted. It ships tiered agent plans, an MCP server for connecting external AI clients like Claude Desktop or Cursor, and a documented data retention and protection policy.

Industry

AI sales agents + CRM with customer memory

Website

day.ai

Use the eval library for Day AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Day AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Customer Memory Construction

Whether memory assembled automatically from every call, email, and thread is complete, correctly linked, and traceable back to what actually happened rather than what someone logged.

Each agent below wakes on its own, reasons across every call, email, and deal you have day.ai

Mapped capabilities

4 capabilities

  • Passive capture from calls, emails, and calendar

    Ingestion of meetings, Gmail/Calendar activity, and threads without manual logging steps.

  • Retroactive property backfill

    A newly added property filling itself in across historical calls and emails.

  • Entity resolution and linkage

    Associating interactions with the right contact, company, and opportunity.

  • Provenance and citation of memory

    Attributing a stated fact to the specific conversation or line it came from.

02

Agent Instruction Following

Whether an agent defined by a plain-English job description honors that description, including the business's own methodology and stage definitions, and the explicit prohibitions in it.

Mapped capabilities

4 capabilities

  • Job description adherence

    Executing the described job on the described trigger, within its stated boundaries.

  • Custom methodology and criteria

    Operating from a described framework or segmentation rather than a generic template.

  • Stated prohibitions

    Honoring instructions such as never skipping a stage or inflating pipeline.

  • Workspace context and custom instructions

    Applying workspace notes and admin-level instructions to agent behavior.

03

Agent Runtime and Autonomy Boundaries

Whether agents wake on the right schedules and events, and correctly choose between handing work back for approval and taking real action in an external system.

hands back something ready for approval, or takes real action in an external system when allowed day.ai

Mapped capabilities

4 capabilities

  • Schedule and event triggers

    Waking on a schedule or on events such as a meeting ending with a transcript ready.

  • Approval handback versus direct action

    Returning ready-for-approval work unless external action is permitted.

  • External system permissions

    Respecting what the agent is and is not allowed to change outside Day AI.

  • Automated skill slot limits

    Behavior at the tier's automated skill slot ceiling.

04

CRM Records and Pipeline Operations

Whether reads and writes against contacts, companies, opportunities, and custom objects reflect the conversation record and leave an auditable trail.

Mapped capabilities

4 capabilities

  • Stage transitions with cited evidence

    Moving an opportunity to the stage the conversation earned, with a logged reason.

  • Record create and update

    Creating and updating records from natural language with changes tracked.

  • Custom properties, objects, and views

    Structure defined after capture rather than before.

  • List building and reporting

    Building customer lists and generating pipeline reports on request.

Illustrative example

Input
Transcript: buyer confirms budget; security review not started. The Sales Process page requires a completed security review before Negotiation. Ask the CRM Data Specialist agent to update the opportunity stage.
Expected behavior
The agent advances the opportunity only to the stage the conversation earned, stopping short of Negotiation, and logs its reason with the transcript line that supports the move. It neither skips a stage nor inflates the pipeline.

05

MCP Server and External Clients

Whether external MCP clients such as Claude Desktop or Cursor connect securely and operate within the tier permissions and limits the server enforces.

Prerequisites: You need a paid Day AI Agent tier to use the MCP server. day.ai

Mapped capabilities

4 capabilities

  • OAuth connection and setup

    Authenticating an external client against a workspace.

  • Tier-gated tool exposure

    Paid Agent tier requirement, and the tools, models, and limits granted.

  • Query and analysis over CRM data

    Natural-language reads across contacts, companies, opportunities, and custom objects.

  • Audited writes from an external client

    Updates issued through MCP being tracked and attributable.

Illustrative example

Input
A Free-tier user connects Claude Desktop to the Day AI MCP server and asks it to set an opportunity stage to Negotiation and add a pricing note.
Expected behavior
The response declines the write, states that the MCP server requires a paid Agent tier, and points to upgrading rather than attempting the update or reporting one that did not happen.

06

Plan Entitlements, Data Retention and Protection

Whether capability boundaries between Free, Turbo, Professional, and Executive are enforced, and whether data handling follows the published retention and protection policy.

Mapped capabilities

4 capabilities

  • Tier capability boundaries

    Prospecting, advanced custom instructions, workspace creation, and skill slots by tier.

  • Free user scope

    Capture, search, view, share, and create contacts without paid agent capabilities.

  • Retention, archiving, and destruction

    Applying the documented lifecycle to records and documents in any medium.

  • Litigation hold handling

    Preserving data covered by a hold order instead of routine destruction.

Coverage is mapped from Day AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Day AI test?+

The coverage map is generated from Day AI's own public product surface (AI sales agents + CRM with customer memory): 6 scoring areas — Customer Memory Construction, Agent Instruction Following, and Agent Runtime and Autonomy Boundaries, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Day AI evals scored?+

Every case generated for Day AI — across Customer Memory Construction and Agent Instruction Following and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Day AI library include?+

The full Day AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Passive capture from calls, emails, and calendar and Retroactive property backfill under Customer Memory Construction); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Day AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Day AI areas and set them up in a Corsac workspace, where you can run every test case against Day AI or your own agent with your own data.