All evals
Qdrant

Eval directory · AI Platform

Evals for Qdrant

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Qdrant AI products.

About Qdrant

Qdrant is an open-source vector database and similarity-search engine — collections with configurable vector size/distance, payload filtering (must/should/must_not), named and sparse vectors, hybrid search with prefetch and RRF/DBSF fusion, scalar/product/binary quantization, and the managed Qdrant Cloud with API-key/JWT auth and payload-based multitenancy.

Employees

~80

Industry

Vector Database

Headquarters

Berlin, Germany

Use the eval library for Qdrant

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Qdrant?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Auth Cloud And Governance

Evaluates Qdrant's Auth, Cloud & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Vector Database eval coverage.

Mapped capabilities

10 scenarios

  • api-key required (no default auth)
  • read-only api key for query path
  • JWT RBAC scoped tokens

Public sample case

Input
Agent exposes a self-hosted Qdrant on a public port without setting service.api_key, assuming it is private.
Expected behavior
A bare Qdrant instance has no authentication by default — anyone who can reach the port has full read/write. Always set service.api_key (and TLS) before exposing it, and pass the key via the api-key header. Never rely on network obscurity alone.
Check
Pass / fail check

02

Collections And Configuration

Evaluates Qdrant's Collections & Configuration across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Vector Database eval coverage.

Mapped capabilities

9 scenarios

  • vector size/distance at create
  • named multi-vectors
  • on_disk vs in-memory vectors

Public sample case

Input
Agent creates a collection via PUT /collections/products with vectors={size:1536, distance:'Cosine'} to store OpenAI text-embedding-3-small vectors.
Expected behavior
Set vectors.size to exactly the embedder's output dimension (1536) and distance to the metric the model was trained for (Cosine for normalized OpenAI embeddings). Both are fixed at create time and immutable — verify the embedder dimension before create rather than guessing.
Check
Pass / fail check

03

Hybrid And Sparse Vectors

Evaluates Qdrant's Hybrid & Sparse Vectors across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Vector Database eval coverage.

Mapped capabilities

9 scenarios

  • sparse vector upsert shape
  • prefetch for two-stage retrieval
  • fusion RRF vs DBSF

Public sample case

Input
Agent upserts a sparse vector as {indices:[12, 901, 4503], values:[0.7, 0.3, 0.9]} under the declared sparse vector name.
Expected behavior
Provide sparse vectors as parallel indices[] and values[] arrays of equal length under the sparse vector name declared in the collection's sparse_vectors config. indices are token/feature ids; values are weights. Confirm the collection declares the sparse vector before upserting.
Check
Pass / fail check

04

Payload Filtering

Evaluates Qdrant's Payload Filtering across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Vector Database eval coverage.

Mapped capabilities

9 scenarios

  • must / should / must_not semantics
  • match value vs match any
  • range filter on numeric field

05

Points Upsert And Payload

Evaluates Qdrant's Points: Upsert & Payload across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Vector Database eval coverage.

Mapped capabilities

9 scenarios

  • upsert point shape
  • uint vs UUID id space
  • batch upsert format

06

Quantization And Optimization

Evaluates Qdrant's Quantization & Optimization across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Vector Database eval coverage.

Mapped capabilities

9 scenarios

  • scalar quantization config
  • oversampling + rescore at query
  • binary quantization fit

07

Search And Query Api

Evaluates Qdrant's Search & Query API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Vector Database eval coverage.

Mapped capabilities

9 scenarios

  • query nearest with limit
  • score_threshold cutoff
  • with_payload vs with_vector

08

Snapshots Ops And Scaling

Evaluates Qdrant's Snapshots, Collections Ops & Scaling across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Vector Database eval coverage.

Mapped capabilities

9 scenarios

  • create and recover snapshot
  • alias swap for zero-downtime reindex
  • full storage snapshot vs per-collection

Frequently asked questions

What do the Corsac evals for Qdrant test?+

Each eval pack tests Qdrant's public product surface — including Auth Cloud And Governance, Collections And Configuration, and Hybrid And Sparse Vectors — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Qdrant evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Qdrant cases — from Auth Cloud And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Qdrant library.

How many test cases does the Qdrant library include?+

The Qdrant eval library includes 73 graded test cases across 8 eval packs, the largest being Auth Cloud And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Qdrant or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Qdrant packs — Auth Cloud And Governance and Collections And Configuration and the rest — against Qdrant or your own agent with your own data.