All evals
11x

Eval directory · Document Agents

Evals for 11x

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for 11x AI products.

About 11x

11x builds autonomous digital workers for revenue workflows, including prospecting, engagement, and inbound-response automation.

Industry

AI / Revenue Automation

Headquarters

San Francisco, CA

Website

www.11x.ai

Use the eval library for 11x

All 26 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Related in Document Agents

All evals →

More Document Agents eval libraries

Coverage map

What would you measure for 11x?

1 area · 26 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Eval Factory Import V1

Evaluates 11x's Eval Factory Import — research & personalization quality, response handling, and lead qualification accuracy — across 26 test cases graded case by case by an LLM judge.

Mapped capabilities

26 scenarios

  • Research & Personalization Quality
  • Response Handling
  • Lead Qualification Accuracy

Example criterion: 11x responses follow required actions, avoid disallowed actions, and maintain risk-aware behavior.

Frequently asked questions

What do the Corsac evals for 11x test?+

Each eval pack tests 11x's public product surface — including Eval Factory Import V1 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the 11x evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the 11x library include?+

The 11x eval library includes 26 graded test cases across 1 eval pack. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against 11x or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against 11x or your own agent with your own data.