All evals
ThoughtSpot

Eval directory · Data Analysis

Evals for ThoughtSpot

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for ThoughtSpot AI products.

About ThoughtSpot

ThoughtSpot is an AI-powered analytics platform that lets anyone explore data and get instant insights using natural language search. Its Sage AI layer translates business questions into SQL, returning charts and answers in seconds without requiring analyst help.

Employees

~1,200

Industry

Business Intelligence

Headquarters

Sunnyvale, CA

Use the eval library for ThoughtSpot

All 3 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Related in Data Analysis

All evals →

More Data Analysis eval libraries

Coverage map

What would you measure for ThoughtSpot?

1 area · 3 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Nl2sql Smoke V1

Evaluates ThoughtSpot's NL2SQL — query intent translation, schema-constrained generation, and result reliability — across 3 test cases graded case by case by an LLM judge.

Mapped capabilities

3 scenarios

  • Query Intent Translation
  • Schema-Constrained Generation
  • Result Reliability

Example criterion: ThoughtSpot reliably translates analyst questions into accurate SQL that returns trustworthy business results.

Frequently asked questions

What do the Corsac evals for ThoughtSpot test?+

Each eval pack tests ThoughtSpot's public product surface — including Nl2sql Smoke V1 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the ThoughtSpot evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the ThoughtSpot library include?+

The ThoughtSpot eval library includes 3 graded test cases across 1 eval pack. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against ThoughtSpot or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against ThoughtSpot or your own agent with your own data.