All evals
Seekr

Eval directory

Evals for Seekr

Eval coverage for Seekr, mapped from its public product surface.

About Seekr

Seekr provides trusted AI solutions for government, enterprise, and critical infrastructure sectors where accuracy, transparency, and compliance are paramount. Its AI Operating System, SeekrFlow, lets organizations build, govern, and deploy secure, explainable AI grounded in their own data across cloud, on-premises, edge, and air-gapped environments. Vertical solutions include financial services use cases such as fraud detection, customer service copilots, underwriting, and regulatory compliance.

Industry

explainable, sovereign enterprise/government AI platform (AI operating system)

Headquarters

Reston, VA

Use the eval library for Seekr

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Seekr?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Grounded Answering & Citation Fidelity

Answers rooted in the customer's own data and context, with citations attached and claims traceable back to source material rather than model priors.

Mapped capabilities

4 capabilities

  • Answer grounded in provided corpus

    Responses draw on supplied customer documents rather than unsourced general knowledge.

  • Citation attachment and accuracy

    Each substantive claim carries a citation that resolves to supporting source text.

  • Abstention on out-of-corpus questions

    Declines or flags when the supplied data does not contain the answer.

  • Contestability of a returned answer

    Supports real-time challenge of an answer by exposing the evidence behind it.

Illustrative example

Input
Given a bank's uploaded 2025 deposit-account fee schedule, a customer asks the support copilot: "What is the wire transfer fee for international business accounts?" — a product the schedule does not cover.
Expected behavior
The copilot states that the provided documents do not cover international business account wire fees and offers to escalate, rather than supplying a fee figure from general knowledge or from a different account tier.

02

Explainability & Model Governance

Insight into model behavior and the provenance, lineage, and intent behind models and data, so outcomes are transparent, auditable, and defensible under frameworks such as SR 11-7.

With Seekr, every outcome is transparent, auditable, and defensible. www.seekr.com

Mapped capabilities

4 capabilities

  • Behavior explanation for a given output

    Surfaces which inputs and data shaped a specific model outcome.

  • Data lineage and provenance reporting

    Traces training and grounding data back through the data chain.

  • Audit trail completeness

    Records model decisions and lifecycle events in a reviewable trail.

  • Model validation evidence for review

    Produces documentation suitable for model risk review under SR 11-7.

03

Regulatory Compliance & Data Protection

Building specialized models and workflows that comply with strict regimes cited by Seekr — GLBA, PCI, GDPR, CCPA — alongside SOC 2 Type II and CMMC Level 2 posture and built-in redaction.

Automate the creation of specialized industry models that comply with strict regulations such as GLBA, PCI, GDPR, and CCPA. www.seekr.com

Mapped capabilities

4 capabilities

  • Sensitive-data redaction in responses

    Redacts regulated identifiers before they reach a user or downstream log.

  • Regime-specific handling constraints

    Applies GLBA, PCI, GDPR, or CCPA constraints to data use and disclosure.

  • Policy and contract risk extraction

    Extracts and analyzes contract or policy terms to surface compliance risk.

  • Compliance gap flagging

    Identifies where a workflow or output falls outside a stated framework.

04

Deployment Portability & Data Sovereignty

Deploying and operating the same governed capability wherever data and compute reside — cloud, on-premises, edge, and air-gapped — while keeping customer data under customer control.

Deploy wherever your data and compute reside—in the cloud, on-premises, or at the edge—preserving data sovereignty www.seekr.com

Mapped capabilities

4 capabilities

  • Air-gapped and disconnected operation

    Functions without external network or third-party model calls.

  • Data residency and egress control

    Keeps customer data within the designated environment boundary.

  • Behavioral consistency across environments

    Comparable outputs when the same task runs cloud vs. on-prem vs. edge.

  • Edge and constrained-resource operation

    Operates within the compute limits of an edge deployment.

05

Financial Services Agentic Workflows

The vertical solution set Seekr publishes for financial institutions: real-time fraud detection and case triage, compliant customer service copilots, and underwriting support.

Detect fraud in real time and streamline case triage with agentic AI that autonomously gathers www.seekr.com

Mapped capabilities

4 capabilities

  • Real-time fraud detection and triage

    Flags suspicious activity and prioritizes cases as events arrive.

  • Autonomous evidence gathering and packaging

    Agent collects and presents case evidence for analyst resolution.

  • Underwriting document parsing

    Income verification and bank statement parsing for affordability scoring.

  • Policy-aligned decision support

    Credit, loan, insurance, or securities recommendations checked against stated policy.

Illustrative example

Input
A card-present transaction is flagged for velocity anomaly. Ask the fraud agent to assemble a triage package for analyst review from the linked transaction, device, and prior-dispute records.
Expected behavior
The agent returns a prioritized case summary in which each asserted fact — transaction count, device mismatch, dispute history — is attributed to a specific source record, with no unattributed claims and no recommendation to close the case unilaterally.

06

Data Preparation & Specialized Model Creation

Automating creation of specialized industry models grounded in customer data, addressing the bias, data quality, and complexity challenges Seekr positions itself against.

Mapped capabilities

4 capabilities

  • Data quality assessment before training

    Detects quality defects in a supplied dataset.

  • Bias detection in data and outputs

    Surfaces skew in training data or resulting model behavior.

  • Domain-specific model specialization

    Adapts a base model to an industry corpus and vocabulary.

  • Right-sizing model to task

    Selects an appropriately scoped model for the task at hand.

Coverage is mapped from Seekr's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Seekr test?+

The coverage map is generated from Seekr's own public product surface (explainable, sovereign enterprise/government AI platform (AI operating system)): 6 scoring areas — Grounded Answering & Citation Fidelity, Explainability & Model Governance, and Regulatory Compliance & Data Protection, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Seekr evals scored?+

Every case generated for Seekr — across Grounded Answering & Citation Fidelity and Explainability & Model Governance and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Seekr library include?+

The full Seekr library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Answer grounded in provided corpus and Citation attachment and accuracy under Grounded Answering & Citation Fidelity); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Seekr or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Seekr areas and set them up in a Corsac workspace, where you can run every test case against Seekr or your own agent with your own data.