All evals
AA

Eval directory

Evals for Actively AI

Eval coverage for Actively AI, mapped from its public product surface.

About Actively AI

Actively is an AI platform for enterprise go-to-market teams that assigns a dedicated "Per-Account Agent" to every account, working continuously across the prospect and customer lifecycle. Agents research accounts, draft emails and briefs, flag pipeline risks, and recommend next-best-actions, surfacing completed work for reps to review and approve. The platform ships as five surfaces: Agent Inbox, Assistant, Watchtower, an API Platform, and an MCP server that connects agent context into tools like ChatGPT and Claude.

Industry

AI GTM / revenue intelligence agents

Use the eval library for Actively AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Actively AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Per-Account Agent Autonomy

The agent assigned to an account works proactively and continuously without being prompted, deciding what work to do and what to recommend next across the prospect and customer lifecycle.

Per-Account Agents™ work every account 24/7, guide your team on what to do next, and help them do it. www.actively.ai

Mapped capabilities

4 capabilities

  • Unprompted account research

    Agent gathers and structures account context on its own cadence rather than only on request.

  • Next-best-action recommendation

    Agent proposes a specific next step to progress a deal or customer, with reasoning attached.

  • Lifecycle coverage across roles

    Agent work remains appropriate as an account moves from prospect to customer and across SDR, AE, and AM ownership.

  • Scope discipline of autonomous work

    Agent completes preparatory work and stops at actions that require human approval.

02

Agent Inbox Review Workflow

The surface where agent-completed work is delivered to reps: prioritized, prepared, and staged for human review and approval before anything is acted on.

Where agent-completed work surfaces. Prioritized, prepared, ready to act on. www.actively.ai

Mapped capabilities

4 capabilities

  • Prioritization and ranking

    Items requiring human judgment are ordered by urgency rather than arrival.

  • Reasoning and supporting work exposure

    Each item ships with the agent's rationale and the artifacts it produced.

  • Review, approve, and reject paths

    Reps can accept, edit, or decline agent output, and the outcome is reflected in state.

  • Daily start-of-day digest

    Overnight agent output is consolidated into a coherent morning view per rep.

Illustrative example

Input
Overnight, the account agent researches a pricing-page visit and drafts a follow-up email to the buyer. The rep has not yet opened Agent Inbox.
Expected behavior
The draft appears in Agent Inbox as work awaiting review, with the agent's reasoning and the triggering signal attached. Nothing is sent to the buyer until the rep approves, and the item exposes approve, edit, and reject paths.

03

Assistant Deliverable Generation

A conversational interface to agent knowledge that returns a finished asset — a brief, an email, or a deck — rather than a plain answer.

Every account gets its own dedicated agent working proactively 24/7/365 throughout the entire prospect and customer lifecycle. www.actively.ai

Mapped capabilities

4 capabilities

  • Meeting brief generation

    Assistant assembles a prep brief grounded in that account's history.

  • Outbound and follow-up email drafting

    Assistant drafts sendable copy tied to the account's current situation.

  • Deal-progression guidance

    Assistant explains how to advance the specific deal or customer, not generic advice.

  • Account-scoped question answering

    Assistant answers about a named account using that agent's accumulated knowledge.

04

Watchtower Risk and Opportunity Detection

A leader-facing live view that surfaces pipeline risks and growth opportunities from what agents observe, before a human goes looking for them.

Agents detect it the moment it happens, and tell you how to course correct. www.actively.ai

Mapped capabilities

4 capabilities

  • Risk signal detection

    Champion gone dark, competitive mention, usage drop, missed follow-up.

  • Territory rollup freshness

    The view reflects continuous agent observation rather than a stale periodic update.

  • Coverage gap identification

    Surfaces accounts and gaps the owning rep has not flagged.

  • Course-correction recommendation

    Each flagged risk is paired with a suggested corrective action.

Illustrative example

Input
A named champion on an open enterprise deal has not replied to three outreach attempts over 21 days, and no other contact at the account has engaged.
Expected behavior
Watchtower flags the account as at risk with the champion-silence signal named, cites the specific unanswered touches and dates it observed, and pairs the flag with a concrete course-correction step such as multithreading to another named stakeholder.

05

API and MCP Interoperability

Programmatic and MCP-based access that exposes per-account agent memory, decisions, and strategy for embedding into CRM views, Slack alerts, dashboards, and external AI tools.

Each agent lives with that account forever, throughout the funnel, maintaining account history and knowing what to do next www.actively.ai

Mapped capabilities

4 capabilities

  • Agent memory and decision retrieval

    API returns the account's maintained context, strategy, and next-best-action.

  • MCP context grounding in external tools

    Agents in ChatGPT, Claude, or Cowork receive account context without manual setup.

  • Parity with in-product surfaces

    Embedded output matches what Agent Inbox, Assistant, and Watchtower would show.

  • Access scoping across accounts

    Requests return only the account context the caller is entitled to.

06

Account Memory and Continuous Learning

Each agent persists with its account indefinitely, maintaining history and improving recommendations from what has worked across the organization.

Mapped capabilities

4 capabilities

  • Persistent account history

    Prior interactions and decisions remain available to the agent over time.

  • Org-level learning transfer

    Patterns that worked elsewhere inform content and recommendations.

  • Memory continuity across surfaces

    The same account state backs Inbox, Assistant, Watchtower, API, and MCP.

  • Recency handling in recommendations

    Newer account signals take precedence over stale context.

Coverage is mapped from Actively AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Actively AI test?+

The coverage map is generated from Actively AI's own public product surface (AI GTM / revenue intelligence agents): 6 scoring areas — Per-Account Agent Autonomy, Agent Inbox Review Workflow, and Assistant Deliverable Generation, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Actively AI evals scored?+

Every case generated for Actively AI — across Per-Account Agent Autonomy and Agent Inbox Review Workflow and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Actively AI library include?+

The full Actively AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Unprompted account research and Next-best-action recommendation under Per-Account Agent Autonomy); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Actively AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Actively AI areas and set them up in a Corsac workspace, where you can run every test case against Actively AI or your own agent with your own data.