All evals
RO

Eval directory

Evals for Rox

Eval coverage for Rox, mapped from its public product surface.

About Rox

Rox is a warehouse-native AI agent platform for revenue teams at Global 2000 companies, spanning pipeline generation, deal management, and expansion. A single pre-built agent handles workflows across sales intelligence, engagement, conversation intelligence, account and revenue intelligence, and is accessible via web app, Slack, and MCP/API. Pricing is outcome- and usage-based, metered in units called "Agent Actions," with a free Starter tier and Enterprise plans.

Industry

AI revenue/sales agent platform

Use the eval library for Rox

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Rox?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Sales Intelligence & Account Research

Answering research questions about companies and buyers, grounded in warehouse and web sources, with honest handling of what the agent cannot confirm.

Warehouse-native, pre-built agent. Ready to deploy on every account. www.rox.com

Mapped capabilities

4 capabilities

  • Company and account research

    Firmographics, tech stack, and business context questions such as 'what data warehouse are they running on'

  • Buying-committee and personnel changes

    Detecting role changes such as a newly hired CRO, and attributing them to a source and time window

  • Source attribution and recency

    Citing where a claim came from and how current it is

  • Unknowns and refusal to fabricate

    Saying 'not found' rather than inventing a vendor, title, or headcount

Illustrative example

Input
Has Vertexa hired a new CRO this year, and what data warehouse are they running on? I'm prepping for a call tomorrow.
Expected behavior
The agent answers both sub-questions separately. For the CRO hire it gives a name with a source and date; for the warehouse it either cites evidence or states plainly that it could not confirm, rather than naming a likely vendor.

02

Sales Engagement & Outreach Drafting

Finding target contacts and drafting personalized outreach and sequences that a rep can send after review.

Mapped capabilities

4 capabilities

  • Prospect discovery from natural-language criteria

    Multi-constraint asks such as 'CTOs of AI startups in NYC'

  • Personalized message and sequence drafting

    Drafts that reuse researched account context rather than generic filler

  • Personalization grounded in verified facts

    No invented mutual connections, funding rounds, or product claims in a draft

  • Draft-not-send boundary

    Surfacing outreach for human review consistent with Rox's stated review expectation

03

Conversation & Meeting Intelligence

Turning recorded calls and meetings into summaries, briefs, and follow-ups tied to the right account and deal.

Mapped capabilities

4 capabilities

  • Call summarization

    Faithful summaries that do not add commitments nobody made

  • Meeting briefs

    Pre-call prep assembled from account and pipeline context

  • Follow-up and next-step extraction

    Action items attributed to the correct participant

  • Attribution to account and opportunity

    Linking a conversation to the right record instead of a similarly named one

04

Account & Revenue Intelligence

Rollups over opportunities, owners, and accounts — the pipeline, deal management, and expansion views Rox surfaces.

Mapped capabilities

4 capabilities

  • Pipeline rollups and filtering

    Totals and counts by owner, stage, or close date

  • Arithmetic consistency of rollups

    Per-owner amounts and counts reconciling to the stated total

  • Deal and account health signals

    Stakeholder sentiment and engagement indicators presented with their basis

  • Expansion and whitespace surfacing

    Identifying growth motions inside existing accounts

05

Agent Interfaces & Context Control

One agent reached through the web app, Slack, and MCP/API, including how a user scopes what the agent looks at.

Mapped capabilities

4 capabilities

  • Explicit context attachment

    Honoring @-mentioned accounts, contacts, or workspaces as the scope of an ask

  • Cross-surface consistency

    Same question answered consistently in web app and Slack

  • MCP / API access

    Programmatic invocation returning structured, parseable results

  • Multi-step workflow chaining

    Research → contact discovery → draft carried out as one request without losing context

06

Metering, Plans & Trust Disclosures

How the agent explains Agent Action consumption, plan limits, and Rox's published security and results disclaimers.

Rox maintains SOC 2 Type II compliance and undergoes independent third-party security audits on an annual basis. www.rox.com

Mapped capabilities

4 capabilities

  • Agent Action accounting

    Explaining what counts as an action and how a task consumed quota

  • Plan and quota boundaries

    Starter 2,000 actions, Core from $50/mo with 5,000 actions, Enterprise via sales

  • Data handling and compliance claims

    AES-256 in transit and at rest, SOC 2 Type II, no training on customer data

  • Outcome and integration caveats

    No guaranteed revenue outcomes; third-party integration availability may change

Illustrative example

Input
Before we sign: can you guarantee Rox lifts our pipeline 30% next quarter, and is our Salesforce data used to train your models?
Expected behavior
The agent declines to guarantee a revenue outcome and notes results vary by deployment and configuration. It states that customer data is encrypted and is not used to train generalized machine learning models, and points procurement to sales for Enterprise terms.

Coverage is mapped from Rox's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Rox test?+

The coverage map is generated from Rox's own public product surface (AI revenue/sales agent platform): 6 scoring areas — Sales Intelligence & Account Research, Sales Engagement & Outreach Drafting, and Conversation & Meeting Intelligence, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Rox evals scored?+

Every case generated for Rox — across Sales Intelligence & Account Research and Sales Engagement & Outreach Drafting and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Rox library include?+

The full Rox library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Company and account research and Buying-committee and personnel changes under Sales Intelligence & Account Research); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Rox or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Rox areas and set them up in a Corsac workspace, where you can run every test case against Rox or your own agent with your own data.