All evals
11x

Eval directory · Document Agents

Evals for 11x

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for 11x AI products.

About 11x

11x builds autonomous digital workers for revenue workflows, including prospecting, engagement, and inbound-response automation.

Industry

AI / Revenue Automation

Headquarters

San Francisco, CA

Website

www.11x.ai

Use the eval library for 11x

All 26 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Related in Document Agents

All evals →

More Document Agents eval libraries

Coverage map

What would you measure for 11x?

1 area · 26 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Eval Factory Import V1

Evaluates 11x's Eval Factory Import — research & personalization quality, response handling, and lead qualification accuracy — across 26 test cases graded case by case by an LLM judge.

Mapped capabilities

26 scenarios

  • Research & Personalization Quality
  • Response Handling
  • Lead Qualification Accuracy

Public sample case

Input
Respond to this scenario: Target prospect: Jane Smith, VP Engineering at TechCorp (Series B SaaS company, 150 employees). Recent LinkedIn activity shows she posted about struggling with engineering team velocity. TechCorp recentl…
Expected behavior
disallowed_actions: fabricate facts, contradict provided scenario constraints · required_actions: state assumptions clearly, reference known context only… · response_style: concis…
Check
Pass / fail check

Public sample case

Input
Respond to this scenario: Target prospect: Mike Johnson, CEO at a 20-person startup. Company website shows they just launched 2 weeks ago. Founder's LinkedIn shows this is his first startup after 10 years at Google. Response qual…
Expected behavior
disallowed_actions: fabricate facts, contradict provided scenario constraints · required_actions: state assumptions clearly, reference known context only… · response_style: concis…
Check
Pass / fail check

Public sample case

Input
Respond to this scenario: Research reveals the target company had layoffs last quarter. Current prospect is the new Head of Sales hired 1 month ago. Response quality rule: Agent should recognize sensitive context (layoffs) and th…
Expected behavior
disallowed_actions: fabricate facts, contradict provided scenario constraints · required_actions: state assumptions clearly, reference known context only… · response_style: concis…
Check
Pass / fail check

Example criterion: 11x responses follow required actions, avoid disallowed actions, and maintain risk-aware behavior.

Frequently asked questions

What do the Corsac evals for 11x test?+

Each eval pack tests 11x's public product surface — including Eval Factory Import V1 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the 11x evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 26 11x cases — from Eval Factory Import V1 (26 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the 11x library.

How many test cases does the 11x library include?+

The 11x eval library includes 26 graded test cases across 1 eval pack, the largest being Eval Factory Import V1 with 26 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against 11x or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 1 11x pack — Eval Factory Import V1 and the rest — against 11x or your own agent with your own data.