All evals
Unstructured

Eval directory · AI Platform

Evals for Unstructured

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Unstructured AI products.

About Unstructured

Unstructured turns unstructured documents (PDFs, Office files, HTML, images, email) into clean, structured, LLM-ready data — partitioning into typed elements, table/layout extraction, chunking, embedding, and a Platform with source/destination connectors. Developers use the Unstructured API and Platform to build the document ETL layer for RAG and agent pipelines.

Employees

~75

Industry

Document ETL

Headquarters

San Francisco, CA

Use the eval library for Unstructured

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Unstructured?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Auth Throughput Pii And Governance

Evaluates Unstructured's Auth, Throughput, PII & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Document ETL for LLMs eval coverage.

Mapped capabilities

10 scenarios

  • API key handling
  • rate limit / 429 backoff
  • async job for throughput

Public sample case

Input
The Unstructured API key is hardcoded in client-side code and logged in request traces.
Expected behavior
Send the API key only in the documented header (api-key / unstructured-api-key) server-side from a secret store; never embed it in client code, URLs, or logs. Rotate keys and scope them per environment.
Check
Pass / fail check

02

Chunking

Evaluates Unstructured's Chunking across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Document ETL for LLMs eval coverage.

Mapped capabilities

9 scenarios

  • by_title section boundaries
  • max_characters hard cap
  • overlap configuration

Public sample case

Input
Operator uses chunking_strategy=by_title so chunks respect section structure, but the source was partitioned with fast and has few real Title elements.
Expected behavior
by_title starts new chunks at Title/section boundaries — it depends on accurate Title detection, which comes from good partitioning (often hi_res). Verify Titles exist before relying on by_title; otherwise chunks degrade toward size-only splits.
Check
Pass / fail check

03

Embedding And Enrichment

Evaluates Unstructured's Embedding & Enrichment across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Document ETL for LLMs eval coverage.

Mapped capabilities

9 scenarios

  • embed runs on chunks not raw docs
  • embedding model dimension match
  • enrichment as metadata not text mutation

Public sample case

Input
An embedding transform is configured before chunking, so a 50-page document is embedded as a single oversized vector.
Expected behavior
Place the embedding transform after partition + chunk so each chunk gets its own vector sized for the embedding model's context. Verify the embedded unit is a chunk (bounded by max_characters), not a whole document.
Check
Pass / fail check

04

Metadata And Element Schema

Evaluates Unstructured's Metadata & Element Schema across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Document ETL for LLMs eval coverage.

Mapped capabilities

9 scenarios

  • page_number preservation
  • filename / filetype metadata
  • parent_id hierarchy

05

Partition Document To Elements

Evaluates Unstructured's Partition (document to elements) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Document ETL for LLMs eval coverage.

Mapped capabilities

9 scenarios

  • element type fidelity
  • element ordering preserved
  • supported file type routing

06

Platform Workflows And Connectors

Evaluates Unstructured's Platform Workflows & Connectors across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Document ETL for LLMs eval coverage.

Mapped capabilities

9 scenarios

  • source-to-destination workflow DAG
  • source connector credentials scope
  • incremental ingest change detection

07

Strategies

Evaluates Unstructured's Strategies across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Document ETL for LLMs eval coverage.

Mapped capabilities

9 scenarios

  • strategy=auto routing
  • fast strategy limits
  • hi_res for complex layout

08

Tables And Layout

Evaluates Unstructured's Tables & Layout across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Document ETL for LLMs eval coverage.

Mapped capabilities

9 scenarios

  • infer_table_structure required
  • text_as_html consumption
  • coordinates / bounding boxes

Frequently asked questions

What do the Corsac evals for Unstructured test?+

Each eval pack tests Unstructured's public product surface — including Auth Throughput Pii And Governance, Chunking, and Embedding And Enrichment — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Unstructured evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Unstructured cases — from Auth Throughput Pii And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Unstructured library.

How many test cases does the Unstructured library include?+

The Unstructured eval library includes 73 graded test cases across 8 eval packs, the largest being Auth Throughput Pii And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Unstructured or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Unstructured packs — Auth Throughput Pii And Governance and Chunking and the rest — against Unstructured or your own agent with your own data.