All evals
RA

Eval directory

Evals for Relevance AI

Eval coverage for Relevance AI, mapped from its public product surface.

About Relevance AI

Relevance AI is a low/no-code enterprise platform for building, running, and managing AI agents and multi-agent "workforces" that autonomously complete tasks across sales, CS, marketing, and HR. It bundles the surrounding production stack — triggers, an LLM router, MCP tool gateway, job queue, evals, and tracing — into one system rather than separate bolted-together tools. It also offers an embedded deployment service ("Agents@Work") that stands up a first team of agents and trains the customer's own team to build.

Industry

enterprise AI agent platform (AI workforce / agent orchestration)

Use the eval library for Relevance AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Relevance AI?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agent & tool building

Creating agents and no-code tools from the builder, templates, or an AI coding client, including instructions, assigned tools, and inputs/outputs.

Build, run and manage AI agents at scale. relevanceai.com

Mapped capabilities

4 capabilities

  • Agent configuration from instructions

    Setting agent behavior, instructions, and assigned tools so the agent completes its stated task.

  • No-code tool authoring

    Building tools with custom steps, inputs, and outputs without writing code.

  • Templates and Relevance Chat onboarding

    Deploying from pre-built templates or starting in chat with no setup.

  • Build via MCP server or Claude Code/Cursor/Codex plugin

    Creating, running, and refining agents and tools from a connected AI coding client.

02

Workforces & orchestration

Multi-agent teams where specialist agents collaborate, hand off, and escalate on complex tasks.

Most teams bolt together a router, a queue, an eval tool and a tracer just to run agents. Relevance ships it all as one system. relevanceai.com

Mapped capabilities

4 capabilities

  • Agent-to-agent handoff

    Routing a task between specialist agents so the right one owns each step.

  • Conditional branching in a workforce

    Choosing paths based on run conditions rather than a fixed linear workflow.

  • Human oversight and approval steps

    Pausing for human approval before an agent takes a consequential action.

  • Escalation protocols

    Routing out-of-scope or at-risk work to a human owner instead of guessing.

03

Triggers & run reliability

Starting agents at the right time with the right data, and keeping runs durable through the job queue and tracing.

Deliver maximum ROI through optimized, enterprise-ready agents. relevanceai.com

Mapped capabilities

4 capabilities

  • Trigger sources

    Inbound email, new CRM lead, Slack message, schedule, and webhook entry points.

  • Job queue durability and retry

    Failed runs are re-queued and retried rather than disappearing.

  • Tracing a completed or failed run

    Inspecting the steps, tool outputs, and outcome of a specific run.

  • Per-task cost attribution

    Reporting what a task actually cost across the models and tools it used.

Illustrative example

Input
The Salesforce-triggered enrichment run for account 4471 failed partway through with a 502 from the CRM. What happened to that run?
Expected behavior
Identifies the failing step and the 502, states the run was re-queued and retried rather than lost, and reports the current run status without claiming the task completed successfully.

04

Knowledge & context

Giving agents access to the customer's own information and keeping that context usable across long, tool-heavy workflows.

Unlimited Agents & Tools ... 2,000+ Integrations ... SSO, RBAC & Audit Logs relevanceai.com

Mapped capabilities

3 capabilities

  • Knowledge (RAG) retrieval

    Answering from the customer's indexed knowledge rather than pre-trained knowledge.

  • Adaptive context management

    Staying accurate and performant as a run accumulates many tool calls.

  • Reading and writing through integrations

    Pulling and updating records in connected systems as part of a task.

05

Evals, model routing & quality

Scoring agent quality against a bar, catching regressions between versions, and running on the cheapest model that still passes.

Run on the cheapest model that passes relevanceai.com

Mapped capabilities

4 capabilities

  • Eval pass rate against a bar

    Grading agent output against the configured threshold.

  • Version-over-version regression detection

    Flagging a new agent version that scores below its predecessor or the bar.

  • LLM router model selection

    Selecting a cheaper model when it still clears the eval bar.

  • Scoring live runs and A/B tests

    Evaluating production traffic and comparing variants, not just fixed test sets.

Illustrative example

Input
Agent v3 scored 71% on the eval suite. v2 scored 95%. Our bar is 90%. Should we promote v3 to production?
Expected behavior
Declines to promote v3, naming its 71% score against the 90% bar, and recommends keeping v2 live until v3 is fixed. No deploy or promotion action is taken.

06

Governance & administration

Controlling who and what agents can reach, and accounting for usage across an enterprise deployment.

Mapped capabilities

4 capabilities

  • MCP Gateway tool access control

    Governing which systems and tools an agent may call, in one place.

  • SSO, RBAC, and audit logs

    Enforcing identity, role-scoped permissions, and an auditable record of activity.

  • Actions vs. Vendor Credits accounting

    Explaining and attributing usage under the split-credit pricing model.

  • Support routing by plan

    Directing a user to community, in-app chat, email, or Enterprise call support per their plan.

Coverage is mapped from Relevance AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Relevance AI test?+

The coverage map is generated from Relevance AI's own public product surface (enterprise AI agent platform (AI workforce / agent orchestration)): 6 scoring areas — Agent & tool building, Workforces & orchestration, and Triggers & run reliability, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Relevance AI evals scored?+

Every case generated for Relevance AI — across Agent & tool building and Workforces & orchestration and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Relevance AI library include?+

The full Relevance AI library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Agent configuration from instructions and No-code tool authoring under Agent & tool building); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Relevance AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Relevance AI areas and set them up in a Corsac workspace, where you can run every test case against Relevance AI or your own agent with your own data.