Exa Labs
For Exa LabsSearch & KnowledgeSearch ApiAnswer Relevance

Highlights Summary Synthesis Grounding

Exa · Exa Labs

Neural web search API — Exa

Evaluates Exa Labs' Highlights Summary & Synthesis Grounding across 12 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Neural web search API eval coverage.

About Exa Labs

Exa (formerly Metaphor) is a neural search API that understands the meaning of queries rather than matching keywords — returning the most relevant URLs and content from the web for any semantic question. Developers use Exa to power research agents, content discovery, and RAG pipelines.

Employees

~30

Industry

Neural Search API

Headquarters

San Francisco, CA

Website

exa.ai

Sample tests· showing 3 of 12

#InputExpected behaviorCheck
01

Patent agent sets highlights with query 'claims directed to transformer attention' on contents block.

Set highlights query to match user intent so returned snippets are query-steered; ground answers in returned highlight text only.

Pass / FailGroundinghigh
02

Deep search with outputSchema for competitor table; ~2s synthesis latency acceptable.

Pass through grounding citations from Exa output object; link claims to result URLs; do not add uncited rows.

Pass / FailGroundingcritical
03

Agent requests structured summary with JSON schema for product pricing fields.

Validate schema conformance and cross-check against source text/highlights; flag [REQUIRES-VERIFICATION] for numbers not directly supported.

Pass / FailGroundinghigh

Unlock full benchmark

9 more test cases

Use this benchmark

How this eval is graded

Grade against expected.ideal_behavior and expected.rubric.

Rubric criteria

  • Exa Labs
  • Search
  • Highlights Summary Synthesis Grounding

Recommended for

ExaExa Labs customers

Works with

Related evals

Frequently asked questions

What does the Highlights Summary Synthesis Grounding eval for Exa Labs Exa test?+

Evaluates Exa Labs' Highlights Summary & Synthesis Grounding across 12 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Neural web search API eval coverage.

How is the Highlights Summary Synthesis Grounding eval scored?+

The judge rubric: Grade against expected.ideal_behavior and expected.rubric.

How many test cases does this eval pack include?+

The Highlights Summary Synthesis Grounding pack for Exa Labs Exa contains 12 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.

How do I run this eval?+

Sign up for Corsac, connect your model or agent endpoint, and run the Highlights Summary Synthesis Grounding pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.

Run this eval in your workspace

Connect your data, configure thresholds, and review results with your team.