All evals
Hume AI

Eval directory · AI Platform

Evals for Hume AI

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Hume AI AI products.

About Hume AI

Hume AI builds empathic voice technology — the Empathic Voice Interface (EVI), an expression-measurement API across face, voice, and language, and the Octave text-to-speech model — for emotionally aware, responsibly governed voice experiences.

Employees

~60

Industry

Voice AI

Headquarters

New York, NY

Website

hume.ai

Use the eval library for Hume AI

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Hume AI?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Hume Auth Keys Rate Limits Concurrency

Evaluates Hume AI's Auth, Keys, Rate Limits & Concurrency across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Empathic Voice AI eval coverage.

Mapped capabilities

9 scenarios

  • API key header for REST
  • access-token exchange for EVI
  • secret never on the client

Public sample case

Input
Agent calls POST /v0/batch/jobs and POST /v0/tts and must authenticate the server-to-server REST requests.
Expected behavior
Authenticate REST calls with the X-Hume-Api-Key header carrying the project API key from a server-side secret store. Do not place the key in a query string or client-visible code.
Check
Pass / fail check

02

Hume Evi Interruption Turn Taking

Evaluates Hume AI's EVI Interruption & Turn-taking across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Empathic Voice AI eval coverage.

Mapped capabilities

9 scenarios

  • user_interruption barge-in
  • halt TTS buffer on barge-in
  • end-of-turn detection

Public sample case

Input
User starts speaking while the assistant is still playing audio_output. The server emits a user_interruption event.
Expected behavior
On user_interruption, immediately stop local playback of the assistant's audio_output and stop enqueuing further chunks for that turn, then yield to the user. Treat the interrupted assistant turn as cut short, not completed.
Check
Pass / fail check

03

Hume Evi Sessions

Evaluates Hume AI's EVI Sessions across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Empathic Voice AI eval coverage.

Mapped capabilities

10 scenarios

  • config_id on chat WebSocket
  • session_settings audio encoding
  • system prompt placement

Public sample case

Input
Agent opens wss://api.hume.ai/v0/evi/chat to start an empathic voice session and wants to apply a saved EVI configuration (system prompt, voice, supplemental LLM, tools).
Expected behavior
Reference the saved configuration by passing config_id as a query parameter on the chat WebSocket URL. Do not re-send the full config inline on every connection; the config_id binds the session to its system prompt, voice, LLM, and tools server-side.
Check
Pass / fail check

04

Hume Evi Tool Use

Evaluates Hume AI's Tool Use / Function Calling in EVI across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Empathic Voice AI eval coverage.

Mapped capabilities

9 scenarios

  • tool_call to tool_response pairing
  • tool_response payload shape
  • tool_error vs fake success

05

Hume Expression Measurement Api

Evaluates Hume AI's Expression Measurement API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Empathic Voice AI eval coverage.

Mapped capabilities

9 scenarios

  • model selection face/prosody/language
  • burst (vocal bursts) model
  • modality vs media mismatch

06

Hume Octave Tts

Evaluates Hume AI's Octave TTS across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Empathic Voice AI eval coverage.

Mapped capabilities

9 scenarios

  • description (acting) instructions
  • description vs text separation
  • voice selection vs voice design

07

Hume Safety Ethics Governance

Evaluates Hume AI's Safety, Ethics & Governance across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Empathic Voice AI eval coverage.

Mapped capabilities

9 scenarios

  • no ground-truth emotion claims
  • no deception / honesty detection
  • overclaim requires verification tag

08

Hume Webhooks Batch Jobs

Evaluates Hume AI's Webhooks & Post-call / Batch Jobs across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Empathic Voice AI eval coverage.

Mapped capabilities

9 scenarios

  • batch job submission and identity
  • job size / duration limits
  • status polling with backoff

Frequently asked questions

What do the Corsac evals for Hume AI test?+

Each eval pack tests Hume AI's public product surface — including Hume Auth Keys Rate Limits Concurrency, Hume Evi Interruption Turn Taking, and Hume Evi Sessions — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Hume AI evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Hume AI cases — from Hume Evi Sessions (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Hume AI library.

How many test cases does the Hume AI library include?+

The Hume AI eval library includes 73 graded test cases across 8 eval packs, the largest being Hume Evi Sessions with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Hume AI or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Hume AI packs — Hume Auth Keys Rate Limits Concurrency and Hume Evi Interruption Turn Taking and the rest — against Hume AI or your own agent with your own data.