All evals
Hightouch

Eval directory · AI Platform

Evals for Hightouch

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Hightouch AI products.

About Hightouch

Hightouch is the composable Customer Data Platform — reverse-ETL from warehouses (Snowflake, BigQuery, Redshift, Databricks) to 200+ SaaS destinations, Customer Studio for visual audience building on top of the warehouse, and AI Decisioning for next-best-action and send-time personalization.

Employees

~150

Industry

Customer Data Platform

Headquarters

San Francisco, CA

Use the eval library for Hightouch

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Hightouch?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ai Decisioning

Evaluates Hightouch's AI Decisioning across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Composable CDP / Reverse ETL eval coverage.

Mapped capabilities

9 scenarios

  • training corpus opt-out
  • treatment vs control experiment integrity
  • action eligibility constraints

Public sample case

Input
Operator enables AI Decisioning across the workspace. A subset of users have previously requested data-processing opt-out under GDPR. Their events still enter the training corpus.
Expected behavior
Filter opt-out users out of the AI Decisioning training corpus at source. Per docs, the training corpus is derived from warehouse data the operator exposes — the operator must enforce opt-out via the Model query. Document the opt-out filter as a contract.
Check
Pass / fail check

02

Audiences Customer Studio

Evaluates Hightouch's Audiences (Customer Studio) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Composable CDP / Reverse ETL eval coverage.

Mapped capabilities

9 scenarios

  • Parent Model identity column
  • trait freshness vs audience refresh
  • audience split A/B mutually exclusive

Public sample case

Input
Operator builds a Customer Studio Parent Model on top of the warehouse but picks an unstable session_id as the identity column instead of customer_id.
Expected behavior
Parent Model identity column must be stable across sessions (customer_id, user_uuid). Per Customer Studio docs, every audience traits + related-events join through this identity — instability fragments users and breaks audience membership. Use identity resolution to merge known aliases.
Check
Pass / fail check

03

Auth Rbac Privacy Governance

Evaluates Hightouch's Auth, RBAC, Privacy & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Composable CDP / Reverse ETL eval coverage.

Mapped capabilities

10 scenarios

  • workspace API key scoping
  • SAML SSO enforcement
  • role separation: Editor vs Viewer vs Admin

Public sample case

Input
Engineer creates a workspace-level API key for a Hightouch automation and checks it into the org's shared monorepo. The key has full write permissions on all syncs.
Expected behavior
Create API keys with the minimum scope needed (per docs, key permissions are configurable) and store them in a secrets manager (Vault, AWS Secrets Manager, Doppler) — never in source control. Rotate on a schedule and revoke on engineer offboarding.
Check
Pass / fail check

04

Destinations

Evaluates Hightouch's Destinations across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Composable CDP / Reverse ETL eval coverage.

Mapped capabilities

9 scenarios

  • Salesforce upsert by external_id
  • HubSpot event vs contact mode
  • Segment write-key per source

05

Models Sql Dbt

Evaluates Hightouch's Models (SQL / dbt) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Composable CDP / Reverse ETL eval coverage.

Mapped capabilities

9 scenarios

  • primary key stability
  • diff strategy: Lightning vs tracking column vs all
  • dbt model reference vs raw SQL

06

Operations Observability Alerts

Evaluates Hightouch's Observability, Run History & Alerts across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Composable CDP / Reverse ETL eval coverage.

Mapped capabilities

9 scenarios

  • alert on sync failure with destination context
  • rejected-row threshold alert
  • run history retention + audit

07

Sources Warehouses

Evaluates Hightouch's Sources (Warehouses) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Composable CDP / Reverse ETL eval coverage.

Mapped capabilities

9 scenarios

  • Snowflake role least privilege
  • BigQuery service-account key handling
  • Databricks Unity Catalog grants

08

Syncs

Evaluates Hightouch's Syncs across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Composable CDP / Reverse ETL eval coverage.

Mapped capabilities

9 scenarios

  • sync mode: upsert semantics
  • sync schedule: cron vs triggered vs streaming
  • partial sync on warning

Frequently asked questions

What do the Corsac evals for Hightouch test?+

Each eval pack tests Hightouch's public product surface — including Ai Decisioning, Audiences Customer Studio, and Auth Rbac Privacy Governance — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Hightouch evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Hightouch cases — from Auth Rbac Privacy Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Hightouch library.

How many test cases does the Hightouch library include?+

The Hightouch eval library includes 73 graded test cases across 8 eval packs, the largest being Auth Rbac Privacy Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Hightouch or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Hightouch packs — Ai Decisioning and Audiences Customer Studio and the rest — against Hightouch or your own agent with your own data.