All evals
xpander.ai

Eval directory

Evals for xpander.ai

Eval coverage for xpander.ai, mapped from its public product surface.

About xpander.ai

xpander.ai is a self-hosted, vendor-neutral platform for building, running, and governing AI agents in enterprises. It pairs Omni, an "Agentic Forward Deployed Engineer" that builds, optimizes, migrates, and fixes agents through chat, with a Platform control plane offering an agent registry, connector/MCP registry, policy engine, runtime profiles, model and budget controls, and observability. It runs on any model and on AWS, GCP, Azure, customer VPCs, Kubernetes, or air-gapped on-prem, with usage-based credit pricing for teams and an annual enterprise license.

Industry

enterprise AI agent platform (build, run, and govern AI agents)

Website

xpander.ai

Use the eval library for xpander.ai

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for xpander.ai?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Omni: Agentic Forward Deployed Engineer

The chat-driven lifecycle assistant that stands in for an FDE — designing agents from a described process, improving them from real run history, and repairing them when the surrounding tooling shifts.

Self-hosted, full-stack platform for building and running agents xpander.ai

Mapped capabilities

4 capabilities

  • Build from a described process

    Turning a natural-language workflow description into an agent with wired tools, instructions, and memory.

  • Optimize from past runs and traces

    Reading prior task outcomes to identify failure patterns and rewrite agent instructions.

  • Repair agents after connector or API changes

    Finding every agent affected by a shifted tool contract and updating it before silent failures.

  • Migrate local agents onto the platform

    Bringing locally cobbled-together agents and flows into governed, consistent cloud execution.

02

Governance & Policy Control Plane

The Platform-side controls operators use to decide who may build, run, approve, publish, and access agents, and which capabilities are sanctioned for production.

xpander is the control plane between agent creation and production. xpander.ai

Mapped capabilities

4 capabilities

  • Agent registry

    Owners, versions, environments, and lifecycle status for every agent in the fleet.

  • Connector & MCP registry

    Approved APIs, tools, MCP servers, and private system integrations available to agents.

  • Policy engine and permission scoping

    Role-based rights to build, run, approve, publish, and access; approval gates on risky steps.

  • Shadow-AI containment

    Bringing agents created outside the platform under sanctioned capability and access control.

03

Deployment & Runtime Portability

Where and how agents actually execute — the self-hosted, vendor-neutral posture across clouds, private networks, and disconnected environments, plus the runtime resources allocated to each agent.

Your compute, your cloud or air-gapped environment, on any model - data and memory stay yours. xpander.ai

Mapped capabilities

4 capabilities

  • Cloud and VPC targets

    Running on AWS, GCP, Azure, or a customer-controlled VPC.

  • Kubernetes and air-gapped on-prem

    Self-deployment onto customer Kubernetes clusters or fully disconnected on-prem environments.

  • Runtime profiles

    CPU, memory, GPU, browser, shell, network, and sandbox settings per agent.

  • Data and memory residency

    Keeping agent data and memory inside customer-controlled compute.

Illustrative example

Input
We're on Azure with a strict no-egress policy. Can agents run entirely inside our own environment, and which plan includes self-deployment on Kubernetes or on-prem?
Expected behavior
Confirms the platform is self-hosted and runs on Azure, customer VPCs, Kubernetes, or air-gapped on-prem, and states that self-deployment on Kubernetes or on-prem, with bring-your-own model keys and SSO/OIDC, falls under the Enterprise annual license rather than the credit-based Team plan.

04

Model Routing, Budget & Credit Accounting

Model-agnostic execution paired with the commercial controls around it: which models are approved, how traffic routes, what it costs in credits, and how spend is capped and attributed.

1 credit per agent wake, 1 credit per tool call, and model tokens at each model’s published rate. xpander.ai

Mapped capabilities

4 capabilities

  • Multi-model support and routing

    Running on frontier, open-weight, or customer fine-tuned models with approved-model routing.

  • Credit metering rules

    One credit per agent wake, one per tool call, model tokens billed at published per-model rates.

  • Spend limits and attribution

    Budget caps, sub-org pooled credits, and per-team usage attribution.

  • Model comparison on real work

    Running the same task across models to compare performance and cost before switching.

Illustrative example

Input
An agent wakes once and makes 6 tool calls during that turn. Ignoring model tokens, what does that cost in credits and dollars on the Team plan?
Expected behavior
Answers 7 credits — one for the wake plus one per tool call — which is $0.07 at one cent per credit, and notes that model tokens are billed separately at each model's published rate.

05

Observability, Audit & Failure Recovery

Visibility into what agents did and the path back from a bad run — traces, tool calls, approvals, costs, and failures, plus the debugging and edge-case work Omni performs against them.

Mapped capabilities

4 capabilities

  • Run traces and tool-call logs

    Step-level records of agent runs, tool invocations, and outcomes.

  • Audit trail of actions and approvals

    Auditable history of every agent action and approval decision.

  • Failed-run debugging

    Diagnosing why a specific run failed and what to change.

  • Cost and failure reporting

    Surfacing per-run cost and failure rates back to operators.

06

Multiplayer Collaboration & Access Surfaces

How people and agents work together — publishing expert-built agents for authorized teams to invoke, and reaching them from the chat and coding surfaces an organization already uses.

Mapped capabilities

4 capabilities

  • Publishing and sharing agents

    Experts publishing agents that authorized teams can invoke without a ticket or meeting.

  • Scoped result sharing

    Delivering agent output only to the permitted group.

  • Chat and IDE surfaces

    Invoking agents from Slack, Teams, WhatsApp, Telegram, Claude Code, or the Omni UI.

  • Unlimited-seat collaboration model

    Adding builders and users without per-seat cost, since pricing follows agent work.

Coverage is mapped from xpander.ai's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for xpander.ai test?+

The coverage map is generated from xpander.ai's own public product surface (enterprise AI agent platform (build, run, and govern AI agents)): 6 scoring areas — Omni: Agentic Forward Deployed Engineer, Governance & Policy Control Plane, and Deployment & Runtime Portability, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the xpander.ai evals scored?+

Every case generated for xpander.ai — across Omni: Agentic Forward Deployed Engineer and Governance & Policy Control Plane and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the xpander.ai library include?+

The full xpander.ai library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Build from a described process and Optimize from past runs and traces under Omni: Agentic Forward Deployed Engineer); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against xpander.ai or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped xpander.ai areas and set them up in a Corsac workspace, where you can run every test case against xpander.ai or your own agent with your own data.