All evals
LlamaIndex

Eval directory · AI Platform

Evals for LlamaIndex

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for LlamaIndex AI products.

About LlamaIndex

LlamaIndex is a data framework for building RAG and agent applications over private data — documents/nodes, indexes (VectorStoreIndex), retrievers and query engines, the IngestionPipeline, plus LlamaParse and LlamaCloud for managed document parsing and retrieval.

Employees

~50

Industry

RAG Framework

Headquarters

San Francisco, CA

Use the eval library for LlamaIndex

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for LlamaIndex?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agents And Workflows

Evaluates LlamaIndex's Agents & Workflows across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's RAG / Data Framework eval coverage.

Mapped capabilities

9 scenarios

  • FunctionTool schema fidelity
  • tool error surfaced to agent
  • FunctionAgent vs ReActAgent choice

Public sample case

Input
A Python function is wrapped as a FunctionTool but its parameters lack type hints and the docstring is empty, so the generated tool schema is untyped and the agent calls it with wrong argument types.
Expected behavior
Give tool functions precise type hints and a clear docstring (or an explicit Pydantic schema / fn_schema) so FunctionTool generates a correct JSON schema the agent can call reliably. Validate arguments against the schema before executing; untyped tools produce malformed calls.
Check
Pass / fail check

02

Documents Nodes And Ingestion

Evaluates LlamaIndex's Documents, Nodes & Ingestion across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's RAG / Data Framework eval coverage.

Mapped capabilities

9 scenarios

  • doc_id stability for upsert
  • excluded metadata keys
  • splitter chunk_size vs context

Public sample case

Input
An IngestionPipeline re-runs nightly over a folder of contracts. The loader assigns a fresh random Document.id_ on every run instead of a stable doc_id derived from the source file.
Expected behavior
Set a stable Document.id_ (e.g. derived from the file path or a content/source key) so the pipeline's docstore can detect unchanged documents and dedup them. With a docstore attached, unchanged docs are skipped and changed docs are upserted — random ids defeat dedup and re-embed everything every ru…
Check
Pass / fail check

03

Embeddings And Vector Stores

Evaluates LlamaIndex's Embeddings & Vector Stores across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's RAG / Data Framework eval coverage.

Mapped capabilities

9 scenarios

  • embedding dimension match
  • query vs text embedding asymmetry
  • embed_batch_size and rate limits

Public sample case

Input
A Pinecone index is created with dimension=1536 (for an older embedder) but the configured embed model now outputs 3072-dim vectors; upserts fail or are silently rejected.
Expected behavior
Ensure the vector store's configured dimension exactly matches the embedding model's output dimension. On an embedder change that alters dimension, create a new collection/index at the right dimension and re-embed — you cannot mix dimensions in one space. Verify dimension before bulk upsert.
Check
Pass / fail check

04

Indexes

Evaluates LlamaIndex's Indexes across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's RAG / Data Framework eval coverage.

Mapped capabilities

9 scenarios

  • StorageContext persist + reload
  • external vector store wiring
  • insert keeps docstore in sync

05

Llamaparse And Llamacloud

Evaluates LlamaIndex's LlamaParse / LlamaCloud across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's RAG / Data Framework eval coverage.

Mapped capabilities

9 scenarios

  • parse mode selection
  • job lifecycle polling
  • result_type markdown vs text

06

Observability Settings And Safety

Evaluates LlamaIndex's Observability, Settings & Safety across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's RAG / Data Framework eval coverage.

Mapped capabilities

10 scenarios

  • Settings global vs local override
  • instrumentation / callback tracing
  • token counting / cost telemetry

07

Retrievers And Query Engines

Evaluates LlamaIndex's Retrievers & Query Engines across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's RAG / Data Framework eval coverage.

Mapped capabilities

9 scenarios

  • similarity_top_k tuning
  • source_nodes for citations
  • CitationQueryEngine attribution

08

Structured Outputs And Extraction

Evaluates LlamaIndex's Structured Outputs & Extraction across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's RAG / Data Framework eval coverage.

Mapped capabilities

9 scenarios

  • output_cls validation handling
  • structured_predict vs raw parsing
  • extraction grounded in source

Frequently asked questions

What do the Corsac evals for LlamaIndex test?+

Each eval pack tests LlamaIndex's public product surface — including Agents And Workflows, Documents Nodes And Ingestion, and Embeddings And Vector Stores — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the LlamaIndex evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 LlamaIndex cases — from Observability Settings And Safety (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the LlamaIndex library.

How many test cases does the LlamaIndex library include?+

The LlamaIndex eval library includes 73 graded test cases across 8 eval packs, the largest being Observability Settings And Safety with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against LlamaIndex or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 LlamaIndex packs — Agents And Workflows and Documents Nodes And Ingestion and the rest — against LlamaIndex or your own agent with your own data.