All evals
A

Eval directory

Evals for AlphaSense

Eval coverage for AlphaSense, mapped from its public product surface.

About AlphaSense

AlphaSense is a market intelligence platform that combines search and generative AI over a curated library of premium financial and business documents, including Tegus expert transcripts, broker research, and company filings. Its Enterprise Intelligence tier extends the same AI search and summarization to a firm's own internal content. Recent additions include Deep Research, an agentic multi-step research mode, and native PowerPoint and Excel assistants for producing deliverables.

Industry

market intelligence and research search platform

Use the eval library for AlphaSense

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for AlphaSense?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Grounded Answering and Citation Fidelity

The core generative search promise: synthesized answers over curated premium content with sentence-level citations and no fabricated support. Covers whether claims trace to real retrieved passages and whether the system declines rather than guesses when the corpus is silent.

generate outputs you can trust, with sentence-level citations and no hallucinations www.alpha-sense.com

Mapped capabilities

4 capabilities

  • Sentence-level citation coverage

    Each asserted claim carries a citation that resolves to a specific passage in a retrieved document.

  • Claim-to-source faithfulness

    Cited passages actually support the entity, figure, direction, and timeframe stated in the answer.

  • Abstention on thin or absent evidence

    System states the corpus does not cover the question instead of producing an unsupported synthesis.

  • Multi-document synthesis and disagreement

    Answers that combine transcripts, broker research, and filings surface conflicting views rather than silently picking one.

Illustrative example

Input
What do expert transcripts from the last two quarters say about customer switching costs in enterprise data platforms? Cite each claim.
Expected behavior
The answer asserts only claims supported by retrieved documents, with each declarative sentence carrying a citation that resolves to a specific passage in an entitled source. Points without retrieved support are omitted or explicitly flagged as uncovered rather than asserted.

02

Deep Research Agentic Runs

The multi-step research mode that plans, searches, iterates, and reasons over thousands of results across a long-running session. Covers plan quality, mid-run adaptation, and whether the final report reflects the work actually performed.

Mapped capabilities

4 capabilities

  • Plan decomposition from a prompt

    A simple or detailed prompt becomes a coherent multi-step research plan with distinguishable sub-questions.

  • Iterative replanning on new findings

    The plan adapts when searching surfaces information that contradicts or extends the initial framing.

  • Long-run progress and interruption behavior

    Behavior across a 10-30 minute run, including visibility into state and handling of cancellation or failure mid-run.

  • Report coverage versus stated plan

    The delivered analysis addresses the plan's sub-questions or explicitly flags what could not be answered.

03

Content Coverage and Retrieval Quality

Retrieval behavior over the curated library of premium financial and business documents, including Tegus expert transcripts, broker and independent research, company filings, news, and financial data. Covers whether the right document types and recency windows are reached.

500+ million premium financial and business documents www.alpha-sense.com

Mapped capabilities

4 capabilities

  • Source-type targeting

    Requests scoped to transcripts, broker research, or filings retrieve from the intended document type.

  • Company and entity resolution

    Tickers, aliases, subsidiaries, and private companies map to the correct entity across sources.

  • Recency and time-window handling

    Queries bounded by period or event return documents from that window, including real-time versus aftermarket research.

  • Retrieval recall on niche topics

    Narrow sector or theme questions surface relevant documents rather than defaulting to well-covered large caps.

04

Enterprise Intelligence and Internal Content

Extension of AI search and summarization to a firm's own content, ingested via API or third-party connections. Covers blended internal and market answers, ingestion behavior, and the confidentiality boundary around firm-hosted data.

24/7 support and educational resources & enterprise-grade data protection www.alpha-sense.com

Mapped capabilities

4 capabilities

  • Blended internal and market answers

    Responses combine firm content with premium sources while making the provenance of each claim distinguishable.

  • Ingestion and connector behavior

    Uploads through the API or third-party connections become searchable with expected metadata and handle failed or partial ingests.

  • Internal-content summarization quality

    Summaries of firm documents stay faithful to the source and do not import outside assumptions.

  • Confidential content isolation

    Firm content stays within its hosting and access boundary and does not leak into unentitled contexts.

05

Deliverable Generation in PowerPoint and Excel

The work-product surface: creating decks and spreadsheet outputs from a research thread, inside the native add-ins or from the platform. Covers scale, template conformance, structured-workspace inputs, and whether citations survive into the artifact.

Mapped capabilities

4 capabilities

  • Deck generation at varying scale

    From a single-slide company profile to a large pitch book, produced from the same research session.

  • Firm template and style-guide conformance

    Uploaded templates or style guides govern fonts, colors, and layouts in the first draft.

  • Generation from structured workspaces

    Slides built from organized inputs such as a VDR workspace reflect that workspace's contents.

  • Citation carry-through into artifacts

    Source attribution survives the transition from research thread to slide or spreadsheet output.

Illustrative example

Input
Using the style guide we uploaded, build a ten-slide company profile on the target from this research thread.
Expected behavior
The assistant returns exactly ten slides sourced from the thread's cited documents, applying the uploaded template's fonts, colors, and layouts on the first draft, and preserves source attribution on slides carrying data rather than emitting default-themed slides.

06

Entitlements, Tiering, and Access Boundaries

Behavior at the commercial and data-protection boundary described on the pricing page: content availability differs between Market Intelligence and Enterprise Intelligence, and access is scoped per seat or enterprise-wide.

Mapped capabilities

3 capabilities

  • Tier-scoped content availability

    Content classes available only in a given tier are not returned or cited for users outside it.

  • Entitlement-aware search results

    Result sets and answer sources reflect the user's licensed sources rather than the full library.

  • Denial messaging on unentitled sources

    Blocked content produces a clear entitlement explanation instead of a silent omission or a vague empty result.

Coverage is mapped from AlphaSense's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for AlphaSense test?+

The coverage map is generated from AlphaSense's own public product surface (market intelligence and research search platform): 6 scoring areas — Grounded Answering and Citation Fidelity, Deep Research Agentic Runs, and Content Coverage and Retrieval Quality, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the AlphaSense evals scored?+

Every case generated for AlphaSense — across Grounded Answering and Citation Fidelity and Deep Research Agentic Runs and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the AlphaSense library include?+

The full AlphaSense library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Sentence-level citation coverage and Claim-to-source faithfulness under Grounded Answering and Citation Fidelity); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against AlphaSense or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped AlphaSense areas and set them up in a Corsac workspace, where you can run every test case against AlphaSense or your own agent with your own data.