All evals
Metabase

Eval directory

Evals for Metabase

Eval coverage for Metabase, mapped from its public product surface.

About Metabase

Metabase is an open source business intelligence platform that lets teams explore data by clicking or writing SQL, with an AI assistant (Metabot) that answers questions guided by a semantic layer of metrics and business logic. It is offered as free open source software or as paid Cloud/self-hosted tiers (Starter, Pro, Enterprise) adding support, embedding, and fine-grained permissions. Queries run against the customer's own database rather than data being ingested into Metabase.

Industry

open source business intelligence and AI analytics

Use the eval library for Metabase

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Metabase?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Metabot AI Question Answering

Natural-language questions answered by the AI assistant, with the query behind every answer available for inspection and follow-up scoping offered rather than assumed.

Inspect the query behind every answer. www.metabase.com

Mapped capabilities

4 capabilities

  • Answering business questions in natural language

    Ranking, growth, and breakdown questions like 'Which plan drives the most revenue?' or 'How much did total accounts grow last year?'

  • Showing the query and data sources behind an answer

    Surfacing the underlying query and the tables or data library entries used, so the answer is verifiable rather than asserted

  • Interpretation beyond the number

    Adding the directional read (share of total, deceleration, channel-wide vs. isolated) without overstating what the data supports

  • Proposing scoped follow-ups

    Offering to re-scope a named metric by time period or dimension instead of silently choosing one

Illustrative example

Input
Which plan drives the most revenue?
Expected behavior
Names the top plan with its revenue figure and share of total, drawn from the governed revenue metric. Identifies the data source used and makes the underlying query inspectable, rather than stating a number with no traceable basis.

02

Semantic Layer & Metric Grounding

Answers guided by the customer's defined metrics and business logic so that an agent and a teammate reach the same numbers.

Mapped capabilities

4 capabilities

  • Resolving questions to defined metrics

    Mapping a phrase like 'total revenue' to the governed metric rather than an ad hoc aggregation

  • Referencing metrics explicitly in responses

    Naming the metric or data library entry used, e.g. @Total Revenue by Plan, @New Accounts by Source

  • Consistency between agent and human paths

    The same metric returning the same result whether reached by Metabot or by clicking through

  • Behavior when no defined metric fits

    Stating the gap or the assumption made instead of inventing a definition

03

Self-Service Exploration & SQL

Exploring connected databases by clicking or by writing SQL, with AI SQL generation, across 20+ supported data sources.

20+ supported data sources www.metabase.com

Mapped capabilities

4 capabilities

  • Point-and-click question building

    Building questions, filters, and summaries without SQL

  • AI SQL generation

    Producing SQL from a natural-language request against the connected schema

  • Questions and dashboards

    Unlimited questions and dashboards as the unit of shared analysis

  • Data source connectivity

    Connecting and querying across the 20+ supported source types

04

Permissions & Data Segregation

Controlling who sees what, down to rows and columns, including multi-tenant isolation and identity-managed group permissions.

Multi tenant data segregation, including row- and column level security and granular permissions www.metabase.com

Mapped capabilities

4 capabilities

  • Row- and column-level permissions

    Data sandboxes filtering table views per group

  • Multi-tenant data segregation

    Isolating customer or team data where boundaries are non-negotiable

  • SSO and group provisioning

    JWT, SAML, or LDAP with pre-mapped groups and automated user provisioning

  • Database-level blocking

    Keeping selected databases entirely off-limits to whole groups

05

Embedded Analytics

Embedding charts and dashboards in a product, including customer-facing AI and brand customization, under the plans that offer it.

Mapped capabilities

4 capabilities

  • Embedding charts and dashboards

    Unlimited embeds with tailored interactivity

  • Customer-facing AI questions

    Letting a product's end customers ask questions inside the embed

  • White-labeling and brand customization

    Fully customizing the embedded experience to the host brand

  • User counting for embedded viewers

    Embed users counting toward seats alongside the internal team

06

Deployment, Plans & Support Fit

Choosing between open source, Cloud, and self-hosted across Starter, Pro, and Enterprise, with the pricing, hosting region, and compliance facts that follow.

Mapped capabilities

4 capabilities

  • Plan and price mapping

    Which tier a requirement lands in, and monthly vs. yearly and per-user pricing

  • Cloud vs. self-hosted trade-offs

    Managed upgrades, backups, and monitoring versus running your own infrastructure

  • Hosting regions and isolation

    US, Europe, Latin America, or Asia-Pacific hosting; single tenancy; air-gapped on-prem

  • Compliance and support expectations

    SOC 2 Type II across tiers, SOC 1 at Enterprise, and the support channel each plan includes

Illustrative example

Input
We self-host and need row- and column-level security to isolate customer data. What is the lowest paid plan that covers it?
Expected behavior
Identifies Pro as the lowest tier including row- and column-level security and multi-tenant data segregation, and notes self-hosting is available at that tier. Does not attribute those permissions to Open Source or Starter.

Coverage is mapped from Metabase's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Metabase test?+

The coverage map is generated from Metabase's own public product surface (open source business intelligence and AI analytics): 6 scoring areas — Metabot AI Question Answering, Semantic Layer & Metric Grounding, and Self-Service Exploration & SQL, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Metabase evals scored?+

Every case generated for Metabase — across Metabot AI Question Answering and Semantic Layer & Metric Grounding and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Metabase library include?+

The full Metabase library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Answering business questions in natural language and Showing the query and data sources behind an answer under Metabot AI Question Answering); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Metabase or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Metabase areas and set them up in a Corsac workspace, where you can run every test case against Metabase or your own agent with your own data.