All evals
Pinecone

Eval directory · AI Platform

Evals for Pinecone

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Pinecone AI products.

About Pinecone

Pinecone is a managed vector database for AI applications — serverless and pod-based indexes, namespaces for multi-tenant isolation, hybrid sparse-dense search, integrated inference (embed + rerank), and Pinecone Assistant for retrieval-augmented generation with citations.

Employees

~150

Industry

Vector Database

Headquarters

New York, NY

Use the eval library for Pinecone

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Pinecone?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Auth Quotas Safety And Governance

Evaluates Pinecone's Auth, Quotas, Safety & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Vector Database eval coverage.

Mapped capabilities

10 scenarios

  • API key in client code
  • rate limit 429 handling
  • PII in metadata

Public sample case

Input
Operator ships a SPA that calls /query directly with PINECONE_API_KEY in browser JavaScript.
Expected behavior
Project-scoped API keys must never reach the browser — they grant full read/write across all namespaces in the project. Proxy through an operator-owned backend that scopes namespace by authenticated user. Rotate any key that touched a client bundle.
Check
Pass / fail check

02

Hybrid Search Sparse Dense

Evaluates Pinecone's Hybrid Search (Sparse-Dense) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Vector Database eval coverage.

Mapped capabilities

9 scenarios

  • sparse_vector shape
  • alpha weighting semantics
  • index sparse-support precondition

Public sample case

Input
Operator builds sparse_vector as a Python dict {token:weight} and upserts to a hybrid-capable index.
Expected behavior
Per docs, sparse_vector is {indices:[int...], values:[float...]} with matching lengths — both lists same order. The dict form must be converted to two parallel arrays. Validate before upsert; mismatched lengths or non-int indices are rejected.
Check
Pass / fail check

03

Index Management

Evaluates Pinecone's Index Management across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Vector Database eval coverage.

Mapped capabilities

9 scenarios

  • serverless vs pod choice
  • distance metric immutability
  • dimension mismatch at upsert

Public sample case

Input
Operator must choose between serverless and pod-based for a workload with spiky read traffic, ~10M vectors at 1536 dim, and unpredictable namespace count.
Expected behavior
Pick serverless via spec.serverless{cloud, region}: auto-scales, pays per read/write/storage unit, supports unbounded namespaces. Pod-based (p1/p2/s1) fits steady QPS with predictable size — do not default to pods just because they look 'production'. Document the dimensional and metric choice (immu…
Check
Pass / fail check

04

Integrated Inference Embed And Rerank

Evaluates Pinecone's Integrated Inference / Embed & Rerank across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Vector Database eval coverage.

Mapped capabilities

9 scenarios

  • embed model dimension alignment
  • input_type=passage vs query
  • rerank top_n trim

05

Namespaces And Multitenant Isolation

Evaluates Pinecone's Namespaces & Multi-tenant Isolation across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Vector Database eval coverage.

Mapped capabilities

9 scenarios

  • namespace per tenant
  • namespace listing
  • namespace deletion atomicity

06

Pinecone Assistant

Evaluates Pinecone's Pinecone Assistant across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Vector Database eval coverage.

Mapped capabilities

9 scenarios

  • file upload + retention
  • chat citations are grounded
  • chat output format

07

Query And Filtering

Evaluates Pinecone's Query & Filtering across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Vector Database eval coverage.

Mapped capabilities

9 scenarios

  • top_k bounds
  • query by id vs vector
  • filter expression DSL

08

Upsert And Updates

Evaluates Pinecone's Upsert & Updates across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Vector Database eval coverage.

Mapped capabilities

9 scenarios

  • batch size <= 100 vectors
  • upsert is replace-by-id
  • update set_metadata semantics

Frequently asked questions

What do the Corsac evals for Pinecone test?+

Each eval pack tests Pinecone's public product surface — including Auth Quotas Safety And Governance, Hybrid Search Sparse Dense, and Index Management — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Pinecone evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Pinecone cases — from Auth Quotas Safety And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Pinecone library.

How many test cases does the Pinecone library include?+

The Pinecone eval library includes 73 graded test cases across 8 eval packs, the largest being Auth Quotas Safety And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Pinecone or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Pinecone packs — Auth Quotas Safety And Governance and Hybrid Search Sparse Dense and the rest — against Pinecone or your own agent with your own data.