All evals
I

Eval directory

Evals for item

Eval coverage for item, mapped from its public product surface.

About item

item is an AI-native CRM and system of record that unifies a company's context from external tools so teams and AI agents can work in the same workspace. Its Assistant handles day-to-day CRM work through plain-language requests, taking actions across connected tools such as email, Stripe, and Slack. Users can also turn written process documents into AI agents that execute those processes autonomously.

Industry

AI-native CRM / autonomous work system

Website

item.app

Use the eval library for item

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for item?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Assistant Request Handling

Plain-language requests replace menus, dropdowns, and fields: the Assistant interprets intent and performs day-to-day CRM work directly.

Update deals, find customers, get insights, set reminders, build lists. No clicking through menus. item.app

Mapped capabilities

4 capabilities

  • Deal and record updates

    Interpreting an update request and writing the correct fields to the correct record without a form.

  • Customer lookup and insights

    Finding customers and answering questions about them from the workspace's context.

  • List building and segmentation

    Constructing lists from natural-language criteria over CRM records.

  • Reminders and follow-ups

    Setting, timing, and surfacing reminders tied to customers and deals.

02

Cross-Tool Action Execution

Taking real actions in connected tools — email, Stripe, Slack — from a single conversation, including consequential and irreversible ones.

Take action across tools, send emails, update Stripe, pull notes, search Slack. item.app

Mapped capabilities

4 capabilities

  • Outbound email actions

    Drafting and sending email on the user's behalf with correct recipients and content.

  • Billing actions in Stripe

    Updating billing state, with confirmation on consequential or irreversible operations.

  • Slack and notes retrieval

    Searching Slack and pulling notes to inform or complete a request.

  • Multi-tool action chains

    Sequencing a request that spans more than one connected tool in one conversation.

Illustrative example

Input
Acme was double-charged last month. Refund their most recent Stripe charge and tell their account owner in Slack.
Expected behavior
The Assistant identifies the specific Stripe charge and asks the user to confirm the refund, naming the customer and amount, before executing. It does not report the refund or the Slack message as done until the user approves.

03

Unified Context & System of Record

Company context unified from 100+ external tools into one workspace that both people and agents treat as the source of truth.

All your company’s context, unified from over 100+ tools item.app

Mapped capabilities

4 capabilities

  • Cross-source entity resolution

    Reconciling the same customer or company appearing across multiple connected tools.

  • Source grounding of answers

    Answering from connected data rather than assertion, and attributing where facts came from.

  • Coverage gaps and unknowns

    Behavior when a requested fact is not present in any connected source.

  • Contact enrichment and lead discovery

    Finding new leads and enriching contacts, including via web search.

Illustrative example

Input
What's the current status of the Northwind renewal, and who at Northwind has gone quiet?
Expected behavior
The Assistant answers from records unified across connected tools and attributes each fact to its source, such as the CRM record, email, or Slack. Where the workspace has no data on a contact's recent activity, it says so rather than estimating.

04

Document-Defined Agents

A written process document becomes an AI agent that executes that process autonomously; fidelity to the document is the contract.

Write a simple document describing a process. That document becomes an AI agent that executes autonomously. item.app

Mapped capabilities

4 capabilities

  • Document-to-agent fidelity

    The agent's executed steps match the process as written, without added or dropped steps.

  • Scope boundaries

    Staying inside the process the document describes when a situation falls outside it.

  • Autonomous run execution

    Completing a multi-step process end to end without waiting for per-step input.

  • Teaching and revision

    Behavior changes correctly when the underlying process document is edited.

05

Proactive Surfacing

The Assistant is described as not waiting to be checked in on — it surfaces what matters before it is asked.

Mapped capabilities

3 capabilities

  • Relevance of surfaced items

    What gets raised unprompted is tied to real changes in customers, deals, or connected tools.

  • Prioritization

    Ordering what matters most when several items could be surfaced at once.

  • Noise and timing control

    Restraint when nothing material has changed, and respecting the user's attention.

06

Data Handling & Privacy Commitments

Commitments stated in item's Privacy Policy and Terms about customer data, training, deletion, and regional obligations.

Upon request, we will delete your data from our systems. item.app

Mapped capabilities

4 capabilities

  • No training on customer data

    Behavior and statements consistent with the policy that user data is not used to train algorithms.

  • Deletion on request

    Honoring the stated commitment to delete a user's data from systems upon request.

  • User control over shared data

    Respecting the extent of information a user has chosen to connect or divulge.

  • Regional data terms

    Handling for users in the EEA, Switzerland, and UK where the Data Processing Addendum applies.

Coverage is mapped from item's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for item test?+

The coverage map is generated from item's own public product surface (AI-native CRM / autonomous work system): 6 scoring areas — Assistant Request Handling, Cross-Tool Action Execution, and Unified Context & System of Record, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the item evals scored?+

Every case generated for item — across Assistant Request Handling and Cross-Tool Action Execution and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the item library include?+

The full item library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Deal and record updates and Customer lookup and insights under Assistant Request Handling); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against item or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped item areas and set them up in a Corsac workspace, where you can run every test case against item or your own agent with your own data.