All evals
turbopuffer

Eval directory · Data Analysis

Evals for turbopuffer

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for turbopuffer AI products.

About turbopuffer

turbopuffer is a serverless vector database designed for high-performance approximate nearest neighbor search at scale. It handles ingest, indexing, and hybrid queries with a simple HTTP API and charges only for storage and queries — no always-on infrastructure.

Employees

~10

Industry

Vector Database

Headquarters

United States

Use the eval library for turbopuffer

All 67 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Related in Data Analysis

All evals →

More Data Analysis eval libraries

Coverage map

What would you measure for turbopuffer?

7 areas · 67 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Cache Warmup Billing And Dr

Evaluates turbopuffer's Cache Warmup Billing & DR across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Serverless vector database eval coverage.

Mapped capabilities

9 scenarios

  • warm-cache preburst
  • billing bytes queried
  • billing bytes returned

Public sample case

Input
p99 latency SLO; cold namespace after weekend idle.
Expected behavior
Call POST /v2/namespaces/flash-sale-vectors/warm before burst; then run ANN queries; region client gcp-us-central1.
Check
Pass / fail check

02

Delete And Tombstone Semantics

Evaluates turbopuffer's Delete & Tombstone Semantics across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Serverless vector database eval coverage.

Mapped capabilities

10 scenarios

  • delete_rows by id
  • delete_by_filter purge
  • Post-delete strong query

Public sample case

Input
RTBF request for three articles; must use write delete_rows.
Expected behavior
POST write with delete_rows for ids; verify with strong consistency query/filter; confirm ids absent.
Check
Pass / fail check

03

Hybrid Query And Filters

Evaluates turbopuffer's Hybrid Query & Filters across 11 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Serverless vector database eval coverage.

Mapped capabilities

11 scenarios

  • ANN rank_by with Eq filter
  • BM25 rank_by
  • Multi-query hybrid RRF

Public sample case

Input
RAG over 8M tickets; need ANN plus attribute filter combined.
Expected behavior
POST query with rank_by ANN, top_k 50, filters using Eq on tier per filter DSL; default strong consistency unless latency tradeoff documented.
Check
Pass / fail check

04

Namespace Isolation And Acl

Evaluates turbopuffer's Namespace Isolation & ACL across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Serverless vector database eval coverage.

Mapped capabilities

9 scenarios

  • Cross-namespace path discipline
  • copy_from_namespace DR
  • Namespace-per-tenant refusal

05

Pagination And Export

Evaluates turbopuffer's Pagination & Export across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Serverless vector database eval coverage.

Mapped capabilities

9 scenarios

  • Export cursor resume
  • include_attributes projection
  • limit vs limit.total

06

Upsert Durability And Wal

Evaluates turbopuffer's Upsert Durability & WAL across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Serverless vector database eval coverage.

Mapped capabilities

10 scenarios

  • Read-after-write strong
  • 512MB batch chunking
  • WAL group commit

07

Vector Schema And Distance

Evaluates turbopuffer's Vector Schema & Distance across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Serverless vector database eval coverage.

Mapped capabilities

9 scenarios

  • Dimension mismatch
  • Distance metric swap
  • ANN vs kNN

Frequently asked questions

What do the Corsac evals for turbopuffer test?+

Each eval pack tests turbopuffer's public product surface — including Cache Warmup Billing And Dr, Delete And Tombstone Semantics, and Hybrid Query And Filters — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the turbopuffer evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 67 turbopuffer cases — from Hybrid Query And Filters (11 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the turbopuffer library.

How many test cases does the turbopuffer library include?+

The turbopuffer eval library includes 67 graded test cases across 7 eval packs, the largest being Hybrid Query And Filters with 11 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against turbopuffer or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 7 turbopuffer packs — Cache Warmup Billing And Dr and Delete And Tombstone Semantics and the rest — against turbopuffer or your own agent with your own data.