Eval directory
Evals for Unstructured
8 evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Unstructured AI products.
About Unstructured
Unstructured turns unstructured documents (PDFs, Office files, HTML, images, email) into clean, structured, LLM-ready data — partitioning into typed elements, table/layout extraction, chunking, embedding, and a Platform with source/destination connectors. Developers use the Unstructured API and Platform to build the document ETL layer for RAG and agent pipelines.
Available eval packs for Unstructured
8 packs ready to run.
Auth Throughput Pii And Governance
PII LeakageUnstructured evals — Auth, Throughput, PII & Governance (relift v3 InfraRed)
Chunking
Unstructured evals — Chunking (relift v3 InfraRed)
Embedding And Enrichment
Unstructured evals — Embedding & Enrichment (relift v3 InfraRed)
Metadata And Element Schema
Unstructured evals — Metadata & Element Schema (relift v3 InfraRed)
Partition Document To Elements
Unstructured evals — Partition (document to elements) (relift v3 InfraRed)
Platform Workflows And Connectors
Unstructured evals — Platform Workflows & Connectors (relift v3 InfraRed)
Strategies
Unstructured evals — Strategies (relift v3 InfraRed)
Tables And Layout
Unstructured evals — Tables & Layout (relift v3 InfraRed)
Why eval Unstructured AI
Unstructured's AI features ship behind brand promises about accuracy, safety, and reliability. Buyers and integrators need to know those promises hold up under adversarial prompts, edge-case workflows, and the long tail of real customer inputs — not just the demo path.
The Corsac eval library for Unstructured measures four dimensions teams care about most when deploying ai platform agents:
- Adversarial robustness — does the agent resist prompt injection, jailbreaks, and social-engineering attempts?
- Workflow quality— does it complete the task buyers were shown in the demo, on inputs that don't look like the demo?
- Safety gates — does it escalate or refuse when it should, and only then?
- Operator quality — does it preserve analyst trust by surfacing the right context at the right time?
Every eval pack above is hand-authored against Unstructured's public product surface and runnable in Corsac with your own data.