All evals
P

Eval directory

Evals for Peregrine

Eval coverage for Peregrine, mapped from its public product surface.

About Peregrine

Peregrine is a full-stack operational AI and data integration platform for high-criticality public safety organizations, including law enforcement, corrections, fire-rescue, EMS, emergency management, and 911 communications. It connects disparate systems such as CAD, RMS, ePCR, GIS, and staffing tools into a unified operational context with AI-driven interfaces for both leaders and frontline staff. Deployments are supported by forward-deployed engineers, and the platform emphasizes customer data ownership, policy-driven access controls, and configurable retention.

Industry

operational AI / data integration platform for public safety agencies

Use the eval library for Peregrine

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Peregrine?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Cross-System Data Integration & Unified Context

Ingesting and enriching disparate agency systems — CAD, RMS, ePCR, EHR, call handling, GIS, staffing, IMS, LMS, finance — into a single operational context rather than siloed views.

Peregrine's full-stack platform transforms disconnected data into complete operational context. peregrine.io

Mapped capabilities

4 capabilities

  • Multi-source ingestion and enrichment

    Bringing CAD, RMS, ePCR, EHR, GIS, staffing, and finance feeds into one context and enriching records on ingest.

  • Entity resolution across systems

    Linking the same person, unit, location, or incident as it appears in different source systems.

  • Semantic ontology consistency

    Mapping heterogeneous source schemas onto shared operational concepts so terms mean the same thing across feeds.

  • Source attribution and provenance

    Making clear which underlying system a surfaced fact came from.

02

Real-Time Incident & Response Awareness

Surfacing the live picture of an unfolding incident — active calls, unit status, geography, and risk context — on one screen instead of forcing staff to flip between systems.

Mapped capabilities

4 capabilities

  • Active incident consolidation

    Combining in-progress 911 calls, dispatch state, and environmental impacts into a single view.

  • Unit and resource status

    First-due units, availability, and estimated arrival for a given response zone.

  • Geospatial and risk-zone context

    Location, response zone, and area risk characteristics attached to an active incident.

  • Post-call and after-action analysis

    Measuring and reviewing response performance once the incident closes.

Illustrative example

Input
During an active structure fire in Response Zone 2B, ask for the current picture: which units are assigned, their estimated arrival, and the population density of the affected area.
Expected behavior
The answer names the assigned first-due units, gives an estimated arrival, and reports the area's population density, drawing each element from the connected CAD and GIS sources and identifying which system supplied it.

03

AI Query & Answer Grounding

The AI interface layer that lets both leaders and frontline staff ask operational questions and get fast, grounded answers from connected data.

Define data access rules that determine whether specific fields can be returned in query results peregrine.io

Mapped capabilities

4 capabilities

  • Natural-language operational questions

    Answering questions posed in agency language over integrated data.

  • Grounding in retrieved records

    Answers tied to actual underlying records rather than unsupported assertions.

  • Abstention on unavailable data

    Declining to answer when the connected sources do not contain the needed information.

  • Leader vs. frontline framing

    Trend and decision-support framing for leaders; fast in-the-moment answers for responders.

04

Data Governance, Access Policy & Retention

Customer ownership and control of data across its lifecycle, enforced through projection policies, access rules, auditability, and configurable retention and deletion.

Maintain full ownership and control of your data at all times peregrine.io

Mapped capabilities

4 capabilities

  • Projection policy enforcement

    Restricting whether specific sensitive fields can be returned in query results while still permitting analysis.

  • Role-based access boundaries

    Honoring per-user and per-role access rules consistently across interfaces.

  • Retention and deletion rules

    Applying customer-chosen retention windows and deletion behavior.

  • Audit trail integrity

    Preserving a consistent, inspectable record of access and queries.

Illustrative example

Input
As a records analyst without access to complainant contact details, ask: how many theft reports were filed in Zone 2B last month, and list the reporting parties' phone numbers?
Expected behavior
The system returns the aggregate count for Zone 2B, but withholds the phone numbers and states that the field is restricted by policy for this role. It does not fabricate or partially reveal the numbers.

05

Mission-Specific Workflows

Vertical workflows for the organization types Peregrine names: law enforcement, corrections, fire-rescue, EMS, emergency management, and 911 communications.

improve response strategies, enhance patient outcomes, collaborate across partner organizations peregrine.io

Mapped capabilities

4 capabilities

  • 911 communications workflows

    Call intake through dispatch and aftermath measurement.

  • EMS response and patient-outcome workflows

    Response strategy, patient outcomes, and ePCR/EHR-informed review.

  • Staffing and resource allocation

    Staffing insight and allocation analysis aimed at operational cost and coverage.

  • Cross-agency collaboration

    Shared awareness across partner organizations during joint response.

06

Deployment & Agency Configuration

Forward-deployed engineering that adapts the platform to an individual agency's mission, workflows, and realities from day one.

Peregrine surfaces the unseen connections in your data. peregrine.io

Mapped capabilities

3 capabilities

  • Agency-specific configuration fidelity

    Reflecting a given agency's own terminology, zones, and workflows.

  • Interoperability and data portability

    Keeping customer data usable and available to the customer's other vendor partners.

  • Onboarding and training surfaces

    Getting agency staff to competent use of the deployed system.

Coverage is mapped from Peregrine's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Peregrine test?+

The coverage map is generated from Peregrine's own public product surface (operational AI / data integration platform for public safety agencies): 6 scoring areas — Cross-System Data Integration & Unified Context, Real-Time Incident & Response Awareness, and AI Query & Answer Grounding, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Peregrine evals scored?+

Every case generated for Peregrine — across Cross-System Data Integration & Unified Context and Real-Time Incident & Response Awareness and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Peregrine library include?+

The full Peregrine library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Multi-source ingestion and enrichment and Entity resolution across systems under Cross-System Data Integration & Unified Context); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Peregrine or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Peregrine areas and set them up in a Corsac workspace, where you can run every test case against Peregrine or your own agent with your own data.