All evals
Elastic

Eval directory · Search & Knowledge

Evals for Elastic

Eval coverage for Elastic, mapped from its public product surface.

About Elastic

Elastic's Elasticsearch Platform combines search, a vector database, and generative AI capabilities to power retrieval for AI applications. It is packaged into three solutions — Elasticsearch (search and RAG), Elastic Observability, and Elastic Security — each available serverless, hosted, or self-managed. Elastic Cloud runs on AWS, GCP, and Azure, with self-managed deployment on-prem or in private cloud.

Industry

search AI platform / vector database with observability and security solutions

Use the eval library for Elastic

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Related in Search & Knowledge

All evals →

More Search & Knowledge eval libraries

Coverage map

What would you measure for Elastic?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Search and Retrieval (Elasticsearch)

Core search, vector database, and generative AI retrieval capabilities used to bring context to AI applications and power RAG.

Mapped capabilities

4 capabilities

  • Vector database and index modes

    VectorDB index mode, Columnar Mode, and auto-calibration as introduced in 9.5; when each applies.

  • RAG and AI application retrieval

    Using Elasticsearch's vector database and AI toolkit to ground answers for AI applications.

  • Multimodal and multilingual embeddings

    Jina AI models covering text, image, audio, and video in one shared embedding space across up to 119 languages, including on-prem/air-gapped operation.

  • Cross-project search

    Querying isolated projects in place for global visibility without moving or duplicating data.

Illustrative example

Input
Can Elasticsearch search images and audio alongside text across many languages? What capability makes that possible, and can it run without internet access?
Expected behavior
Points to Jina AI models, which place text, images, audio, and video in one shared embedding space across up to 119 languages, and notes those models now run on-prem for air-gapped environments.

02

Elastic Observability

Unified, open observability for apps and infrastructure, powered by machine learning and analytics, for faster problem resolution.

Mapped capabilities

4 capabilities

  • Logs and metrics workloads

    Positioning and workload fit for logs and metrics, including the stated performance and cost comparisons against Prometheus and Datadog.

  • OpenTelemetry ingest (EDOT)

    Elastic Distributions of OpenTelemetry for instrumenting and shipping telemetry.

  • Kubernetes observability via MCP

    The MCP app for Elastic Observability that surfaces Kubernetes agent skills inside an AI chat interface.

  • Problem resolution workflow

    Moving from signal to resolution using unified observability data and ML-driven analysis.

03

Elastic Security Operations

AI-driven security analytics spanning unified SIEM and XDR for detection, investigation, and response.

Elastic Workflows, now generally available in 9.4, brings native automation to Elastic Security www.elastic.co

Mapped capabilities

4 capabilities

  • Detection and AI-driven alert triage

    Alert triage capabilities delivered in 9.5 and the agentic SecOps positioning.

  • Elastic Workflows automation

    Native scripted playbooks generally available in 9.4, replacing a bolt-on SOAR and its integration overhead.

  • Agentic investigation

    AI agents reasoning through investigations with direct access to alerts, cases, and investigation data.

  • Unified SIEM and XDR scope

    What the platform consolidates natively versus what previously required separate tooling.

04

Deployment Models and Operations

Choosing among serverless, hosted, and self-managed, and the operational control, scaling, and locality trade-offs each implies.

Jina AI models now run on‑prem for air‑gapped environments www.elastic.co

Mapped capabilities

4 capabilities

  • Serverless vs hosted vs self-managed fit

    Control over hardware, cluster size, node count, and versions versus fully managed simplicity.

  • Scaling behavior

    Automatic scaling on search and indexing load versus custom cluster capacity control and storage-based autoscaling in ECE/ECK.

  • Cloud and region coverage

    AWS, GCP, Azure, Alibaba, and FedRAMP for hosted; 60 regions across AWS, Azure, and GCP for serverless.

  • On-prem, private cloud, and air-gapped

    Deploying anywhere on-prem or in private cloud, including air-gapped environments.

Illustrative example

Input
We need Elastic in an air-gapped on-prem environment with full control over cluster size and versions. Which deployment option fits, and can serverless work for us?
Expected behavior
Recommends self-managed for on-prem or private cloud, where the customer controls deployment location, hardware, orchestration, cluster size, node count, and versions. Notes serverless is fully managed by Elastic and generally available only on AWS, GCP, and Azure, so it cannot serve an air-gapped site.

05

Pricing and Packaging

How each deployment option is billed and which capabilities are available where, for evaluation and budgeting decisions.

Mapped capabilities

4 capabilities

  • Billing model by deployment

    Resource-based (hosted), usage-based (serverless), and license-based on nodes and used RAM (self-managed).

  • Payment terms

    Pay-as-you-go monthly versus prepaid commitments.

  • Capability availability by tier

    All solutions and platform capabilities on hosted versus most solution and platform capabilities on serverless.

  • Solution-level packaging

    Elasticsearch, Observability, and Security each offered serverless, hosted, or self-managed.

06

Developer Docs, APIs, and Agent Integration

The surfaces developers and AI coding agents use to build on, operate, and upgrade the platform.

Mapped capabilities

4 capabilities

  • API references

    Elasticsearch, Kibana, and Elastic Cloud API documentation.

  • Client libraries

    Official clients for Java, .NET, Python, and others.

  • Elastic skills for AI agents

    Official skills that teach AI coding agents to work with Elasticsearch, Kibana, Fleet, and the rest of the stack.

  • Versioning and upgrades

    Docs covering Elastic Stack 9.0+ (latest 9.5.0), Serverless, previous versions, and the upgrade guide.

Coverage is mapped from Elastic's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Elastic test?+

The coverage map is generated from Elastic's own public product surface (search AI platform / vector database with observability and security solutions): 6 scoring areas — Search and Retrieval (Elasticsearch), Elastic Observability, and Elastic Security Operations, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Elastic evals scored?+

Every case generated for Elastic — across Search and Retrieval (Elasticsearch) and Elastic Observability and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Elastic library include?+

The full Elastic library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Vector database and index modes and RAG and AI application retrieval under Search and Retrieval (Elasticsearch)); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Elastic or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Elastic areas and set them up in a Corsac workspace, where you can run every test case against Elastic or your own agent with your own data.