All evals
Databricks

Eval directory · Data Analysis

Evals for Databricks

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Databricks AI products.

About Databricks

Databricks is the Data + AI Company, providing a unified lakehouse platform for data engineering, analytics, and machine learning at scale. Thousands of organizations use Databricks to build, train, and deploy AI and ML workloads on their own data.

Employees

~6,000

Industry

Data & AI Platform

Headquarters

San Francisco, CA

Use the eval library for Databricks

All 3 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Related in Data Analysis

All evals →

More Data Analysis eval libraries

Coverage map

What would you measure for Databricks?

1 area · 3 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Assistant Nl2sql Smoke V1

Evaluates Databricks' Assistant NL2SQL — query intent translation, schema-constrained generation, and result reliability — across 3 test cases graded case by case by an LLM judge.

Mapped capabilities

3 scenarios

  • Query Intent Translation
  • Schema-Constrained Generation
  • Result Reliability

Public sample case

Input
How many orders are completed?
Expected behavior
sql: SELECT COUNT(*) FROM orders WHERE status = 'completed';
Check
Pass / fail check

Public sample case

Input
Total revenue from completed orders?
Expected behavior
sql: SELECT SUM(amount) FROM orders WHERE status = 'completed';
Check
Pass / fail check

Public sample case

Input
Number of customers in each country?
Expected behavior
sql: SELECT country, COUNT(*) FROM customers GROUP BY country;
Check
Pass / fail check

Example criterion: Databricks reliably translates analyst questions into accurate SQL that returns trustworthy business results.

Frequently asked questions

What do the Corsac evals for Databricks test?+

Each eval pack tests Databricks's public product surface — including Assistant Nl2sql Smoke V1 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Databricks evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 3 Databricks cases — from Assistant Nl2sql Smoke V1 (3 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Databricks library.

How many test cases does the Databricks library include?+

The Databricks eval library includes 3 graded test cases across 1 eval pack, the largest being Assistant Nl2sql Smoke V1 with 3 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Databricks or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 1 Databricks pack — Assistant Nl2sql Smoke V1 and the rest — against Databricks or your own agent with your own data.