All evals
Hebbia

Eval directory · Search & Knowledge

Evals for Hebbia

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Hebbia AI products.

About Hebbia

Hebbia is an AI platform that enables knowledge workers — primarily in finance and law — to perform complex research and analysis over large corpora of documents. Its retrieval and synthesis capabilities go beyond keyword search to reason across entire document sets.

Employees

~100

Industry

AI Research & Knowledge Management

Headquarters

New York, NY

Website

hebbia.ai

Use the eval library for Hebbia

All 28 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Related in Search & Knowledge

All evals →

More Search & Knowledge eval libraries

Coverage map

What would you measure for Hebbia?

2 areas · 28 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Grounding Smoke V1

Smoke test for enterprise search & Q&A grounding and citations.

Mapped capabilities

3 scenarios

  • Grounded Answer Generation
  • Citation Precision
  • Hallucination Resistance

Example criterion: Hebbia delivers grounded, citation-faithful answers with strong resistance to unsupported hallucinations.

02

Eval Factory Import V1

Evaluates Hebbia's Eval Factory Import — extraction accuracy, citation accuracy, and reasoning quality — across 25 test cases graded case by case by an LLM judge.

Mapped capabilities

25 scenarios

  • Extraction Accuracy
  • Citation Accuracy
  • Reasoning Quality

Example criterion: Hebbia responses follow required actions, avoid disallowed actions, and maintain risk-aware behavior.

Frequently asked questions

What do the Corsac evals for Hebbia test?+

Each eval pack tests Hebbia's public product surface — including Grounding Smoke V1 and Eval Factory Import V1 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Hebbia evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 28 Hebbia cases — from Eval Factory Import V1 (25 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Hebbia library.

How many test cases does the Hebbia library include?+

The Hebbia eval library includes 28 graded test cases across 2 eval packs, the largest being Eval Factory Import V1 with 25 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Hebbia or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 2 Hebbia packs — Grounding Smoke V1 and Eval Factory Import V1 and the rest — against Hebbia or your own agent with your own data.