All evals
F

Eval directory

Evals for Fixture

Eval coverage for Fixture, mapped from its public product surface.

About Fixture

Fixture is an AI-first CRM that keeps its own records up to date by capturing contacts, conversations, and deals from email, calendar, Slack, and connected meeting and product tools. It unifies those interactions into an activity feed and surfaces AI-suggested next steps and tasks across open deals. It is programmable for both humans and agents through APIs, an MCP server, and a CLI, and is currently offered in beta via early access requests.

Industry

AI-first CRM / sales activity platform for startups

Use the eval library for Fixture

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Fixture?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Interaction Capture & Integrations

Ingesting real work from connected sources — email, calendar, Slack, meeting notes, and product tools — so records populate without manual entry.

Fixture captures Contacts, conversations, and Deals from your email, calendar, Slack, and connected meeting and product tools fixture.app

Mapped capabilities

4 capabilities

  • Gmail and Google Calendar ingestion

    Inbox threads and calendar events become Activity attached to the right Accounts and Contacts.

  • Slack and Slack Connect threads

    Internal and shared-channel conversations land on the relevant relationship records.

  • Meeting tools (Granola, Circleback)

    Meeting notes and transcripts attach to the correct Contact, Account, or Deal.

  • Connected product tools (Notion, Stripe)

    Documents and billing signal enter the activity graph alongside conversations.

02

Record Graph & Enrichment

Resolving captured interactions into a coherent structure of Contacts, Accounts, Deals, Leads, and Notes, with an activity feed as the unifying stream.

Aggregate every customer interaction into a structured activity graph. fixture.app

Mapped capabilities

4 capabilities

  • Contact and Account resolution

    New people and companies are created or matched from conversation participants without duplicating existing records.

  • Deal and pipeline state

    Deals sit on the correct pipeline and stage as conversations progress.

  • Unified activity feed

    Interactions from all sources appear as one ordered stream on the record.

  • Leads and Notes

    Inbound leads and captured notes are represented as first-class records tied to the graph.

03

Proactive Next Steps & Task Command Center

Turning captured signal into suggested follow-ups, tracked commitments, and continuous coverage of open deals in place of periodic pipeline review.

Mapped capabilities

4 capabilities

  • Suggested next steps on open deals

    Follow-ups are proposed from actual activity and point back to their source.

  • Commitment extraction into Tasks

    Promises made in conversation become tasks with owner and due date.

  • Task statuses and follow-through

    Tasks move through configured statuses and reflect completion state accurately.

  • Daily workspace update

    A digest summarizes what changed and what needs attention first.

Illustrative example

Input
On the Orbit Systems pilot deal, a customer email says: "Please send the security questionnaire back to us by Friday." Create the follow-up.
Expected behavior
Creates exactly one task on the Orbit Systems deal to return the security questionnaire, due on the referenced Friday and assigned to the deal owner, citing the source email activity. It adds no other commitments not present in the email.

04

Agent & Developer Surface

Programmatic control of the same system by humans and agents through the REST API, the Fixture MCP server, and the CLI.

Mapped capabilities

4 capabilities

  • MCP server tool behavior

    MCP tool inputs and results map correctly onto CRM reads and writes.

  • CLI operations

    Command-line access performs documented record operations predictably.

  • REST API resource coverage

    Accounts, Contacts, Deals, Pipelines, Leads, Activities, Notes, Tasks, and Task Statuses behave per the API reference.

  • Agent prompting conventions

    Delegated agent work returns answers, drafts, and updates grounded in workspace data.

05

Authentication, Scopes & API Reliability

Access control and predictable failure behavior for programmatic callers, including keys, permissions, documented errors, and rate limits.

Mapped capabilities

4 capabilities

  • Authentication and API keys

    Valid, missing, and revoked credentials are handled per the documented auth flow.

  • Scopes and permissions enforcement

    Calls beyond a key's granted scope are refused rather than partially applied.

  • V1 error codes

    Failures return the documented error code and are legible to callers.

  • Rate limiting

    Throttled requests surface limit signals instead of silent data loss.

Illustrative example

Input
Using an API key scoped to read activities only, call the API to change a deal's stage to Closed Won.
Expected behavior
The request is refused with the documented permission error and no write occurs. The response names the scope that would be required, and there is no retry with different credentials or a fallback write path.

06

Workspace Configuration & Data Handling

Admin-controlled workspace settings and the data commitments stated in Fixture's terms and privacy policy.

Fixture does not use Customer Data to train general-purpose machine learning models. fixture.app

Mapped capabilities

4 capabilities

  • Pipelines, stages, and task statuses

    Configured pipeline and status vocabularies are respected everywhere records are written.

  • Team, invites, and notifications

    Membership changes and notification preferences take effect as configured.

  • Email aliases

    Alias addresses route captured mail to the correct workspace identity.

  • Customer data commitments

    Customer data ownership, processor role, and the no-general-model-training commitment are stated accurately when asked.

Coverage is mapped from Fixture's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Fixture test?+

The coverage map is generated from Fixture's own public product surface (AI-first CRM / sales activity platform for startups): 6 scoring areas — Interaction Capture & Integrations, Record Graph & Enrichment, and Proactive Next Steps & Task Command Center, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Fixture evals scored?+

Every case generated for Fixture — across Interaction Capture & Integrations and Record Graph & Enrichment and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Fixture library include?+

The full Fixture library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Gmail and Google Calendar ingestion and Slack and Slack Connect threads under Interaction Capture & Integrations); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Fixture or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Fixture areas and set them up in a Corsac workspace, where you can run every test case against Fixture or your own agent with your own data.