All evals
R

Eval directory

Evals for Rescript

Eval coverage for Rescript, mapped from its public product surface.

About Rescript

Rescript is an AI platform for legal, compliance, and public policy teams that researches, monitors, and organizes laws and policy across federal, state, and local jurisdictions. It combines natural-language research grounded in cited sources, configurable alerts on legislation, hearings, rules, and comment windows, and work-product generation such as 50-state surveys, memos, and reports. It also covers hearings and meetings, transcribing proceedings and producing post-hearing memos with speaker identification.

Industry

legal & public policy intelligence AI

Headquarters

Washington, DC

Use the eval library for Rescript

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Rescript?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Grounded Research & Citation Fidelity

Natural-language questions about current law answered from the platform's source set, with every claim traceable to a document the user can open.

helping enterprises, law firms, and public affairs teams research and track policy across federal, state, and local jurisdictions www.rescript.ai

Mapped capabilities

4 capabilities

  • Natural-language legal question answering

    Answers substantive questions (e.g. what regulations govern PBM transparency) using retrieved statutes, regulations, and other in-scope material.

  • Citation attachment and resolution

    Each asserted legal proposition carries a citation that resolves to a real, retrieved source rather than a paraphrase or unsupported summary.

  • Source-type selection across the eight source families

    Chooses among Web, Statutes, Regulations, Executive Orders, Legislation, Proposed Regulations, Public Comments, and Hearings appropriately for the question asked.

  • Grounding on user-uploaded files

    Incorporates documents added by the user alongside platform sources, and distinguishes user-supplied material from platform sources in the answer.

Illustrative example

Input
What regulations govern PBM transparency in Texas? Limit the answer to provisions currently in effect, and cite the specific sections.
Expected behavior
Answers using Texas statutes and regulations retrieved from the source set, attaching a resolvable citation to each legal assertion. Federal or other-state material is not presented as Texas law, and gaps are stated rather than filled in.

02

Jurisdictional Coverage & Multi-State Surveys

Handling of federal, state, and local jurisdictions, including the 50-state survey work product and the disambiguation problems that come with it.

Produce and share 50-state surveys, memos, and reports with stakeholders. www.rescript.ai

Mapped capabilities

4 capabilities

  • Jurisdiction scoping and filtering

    Restricts answers to the jurisdiction(s) the user asked about across federal, state, and local levels.

  • 50-state survey generation

    Produces a per-jurisdiction survey with consistent columns/questions applied across states.

  • Cross-jurisdiction conflict and divergence handling

    Represents differences between jurisdictions accurately rather than generalizing one state's rule to others.

  • Coverage gaps and no-law-on-point states

    States explicitly when a jurisdiction has no applicable provision instead of filling the cell with an inferred answer.

03

Monitoring, Alerts & Watchlists

Configurable tracking of policy movement from first signal through proposal and adoption, organized by client or issue area.

Configure and receive alerts for legislation, hearings, rules, and comment windows by client or issue area. www.rescript.ai

Mapped capabilities

4 capabilities

  • Alert configuration by client or issue area

    Sets up and scopes tracking rules so alerts route to the right matter, client, or topic.

  • Event-type coverage

    Detects and reports legislation changes, hearing postings, rule activity, and comment windows as distinct event types.

  • Comment-window and deadline tracking

    Surfaces open comment periods and their due dates while the window is still actionable.

  • Watchlist maintenance and lifecycle state

    Keeps tracked items current as a measure moves from introduction to proposal to adoption.

04

Hearings & Meetings Coverage

Live and post-hoc handling of hearings, panel meetings, and markups — transcription, speaker identification, and the resulting memo.

turns the record into a concise post-hearing memo with best-in-class speaker identification www.rescript.ai

Mapped capabilities

4 capabilities

  • Transcription of proceedings

    Captures the record of hearings, panel meetings, markups, and livestreams.

  • Speaker identification and attribution

    Attributes statements to the correct named participant and marks segments it cannot attribute.

  • Post-hearing memo generation

    Condenses a proceeding into a memo covering the substantive exchanges the user cares about.

  • Follow-up questions against the record

    Answers subsequent questions using the hearing transcript as the grounding source.

Illustrative example

Input
Cover today's Senate committee hearing and send me a post-hearing memo on what each witness said about drug pricing.
Expected behavior
Produces a memo attributing statements to named participants, drawing only on drug-pricing exchanges in the transcript. Segments it cannot confidently attribute are marked as unidentified rather than assigned to a speaker.

05

Work Product Generation & Delivery

Turning research and monitoring output into memos, reports, comment summaries, and shareable artifacts for stakeholders.

Turn complexity into clarity with an AI system that can research, track, and organize laws and public policy www.rescript.ai

Mapped capabilities

4 capabilities

  • Memo and report drafting

    Produces client- or stakeholder-ready written work product from platform material.

  • Public comment summarization

    Condenses a body of submitted public comments into themes and positions.

  • Format and structure adherence

    Follows the requested work-product shape, length, and section structure.

  • Sharing with stakeholders

    Delivers finished work product to the intended recipients from within the platform.

06

Reliability, Traceability & Review

The stated commitment that outputs are reviewable — answers, updates, and citations trace back to their source material, and uncertainty is surfaced rather than smoothed over.

Get answers grounded in documents, citations, and jurisdictional context. www.rescript.ai

Mapped capabilities

4 capabilities

  • Traceback from claim to source document

    Every answer, update, and citation can be followed to the underlying material behind it.

  • Fabrication resistance

    Does not produce citations, bill numbers, or provisions absent from the retrieved source set.

  • Uncertainty and abstention

    Says the law is unsettled, unavailable, or outside coverage instead of asserting an answer.

  • Currency and as-of dating

    Makes clear the point in time an answer reflects, given that tracked law changes.

Coverage is mapped from Rescript's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Rescript test?+

The coverage map is generated from Rescript's own public product surface (legal & public policy intelligence AI): 6 scoring areas — Grounded Research & Citation Fidelity, Jurisdictional Coverage & Multi-State Surveys, and Monitoring, Alerts & Watchlists, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Rescript evals scored?+

Every case generated for Rescript — across Grounded Research & Citation Fidelity and Jurisdictional Coverage & Multi-State Surveys and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Rescript library include?+

The full Rescript library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Natural-language legal question answering and Citation attachment and resolution under Grounded Research & Citation Fidelity); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Rescript or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Rescript areas and set them up in a Corsac workspace, where you can run every test case against Rescript or your own agent with your own data.