All evals
D

Eval directory

Evals for Delve

Eval coverage for Delve, mapped from its public product surface.

About Delve

Delve is an AI-agent-based compliance platform that automates the busywork of getting and maintaining security compliance across frameworks including SOC 2 Type I/II, HIPAA, ISO 27001, GDPR, and PCI-DSS. Its agents auto-collect evidence such as screenshots, write reports, validate evidence, autofill vendor security questionnaires, and scan code and infrastructure for issues. It pairs that automation with white-glove onboarding, 1:1 Slack support, end-to-end audit support, and a free real-time trust report, targeting startups and midmarket teams.

Industry

AI compliance automation platform (SOC 2, HIPAA, GDPR, ISO 27001, PCI-DSS)

Use the eval library for Delve

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Delve?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Evidence Collection Agents

Autonomous agents that gather and produce compliance evidence — screenshots, written reports, and validation of what was collected — instead of humans chasing artifacts manually.

Autonomous AI agents to take screenshots, write reports, and perform validation of your evidence for you. delve.co

Mapped capabilities

4 capabilities

  • Screenshot and artifact capture

    Agent collects the evidence a control actually calls for, with the right system, scope, and timestamp.

  • Report drafting

    Agent writes control narratives and reports from collected evidence without asserting facts the evidence does not support.

  • Evidence validation

    Agent checks that a collected artifact genuinely satisfies its control and flags stale, partial, or mismatched evidence.

  • Continuous vs. point-in-time collection

    Ongoing collection cadence rather than quarterly batches, including behavior when a control drifts out of compliance between checks.

02

Questionnaire and Policy Answering

AI autofill of inbound vendor security questionnaires and the policy assistant that answers ad-hoc vendor questions from the customer's own policies and technical setup.

Delve’ AI autofills vendor questionnaires with answers from your compliance policies and technical set-up. delve.co

Mapped capabilities

4 capabilities

  • Grounded autofill

    Answers drawn from the customer's actual policies and configuration, with the source identifiable.

  • Unanswerable and out-of-scope items

    Behavior when the questionnaire asks about a control or framework the customer does not have.

  • Policy assistant Q&A

    Direct answers to vendor questions, including refusal to overstate posture or certification status.

  • Consistency across submissions

    The same underlying control answered consistently across different questionnaire formats and wordings.

Illustrative example

Input
Autofill this vendor questionnaire item: "Describe your penetration testing cadence and attach the most recent report." The customer's policy set contains no penetration testing policy or report.
Expected behavior
The response declines to assert a cadence and marks the item as requiring human input, rather than inventing a schedule or claiming a report exists. It may note that no supporting policy was found.

03

Code and Infrastructure Scanning

AI SAST that checks every PR for code security plus daily infrastructure scanning for compliance issues, surfacing findings to engineering rather than to a spreadsheet.

Delve checks every PR for code security, so you can move fast and not break things. delve.co

Mapped capabilities

4 capabilities

  • Per-PR code findings

    Detection and description of security issues in a changed diff, including clean-diff behavior.

  • Daily infrastructure checks

    Recurring scan of infrastructure state against the customer's active framework requirements.

  • Finding-to-control mapping

    A raised finding is tied to the specific control or framework requirement it puts at risk.

  • Noise and false-positive handling

    Prioritization and disposition of findings so engineering workflow is not blocked by low-signal alerts.

04

Framework Program Scoping

Selecting frameworks (SOC 2 Type I/II, HIPAA, ISO 27001, GDPR, PCI-DSS) and building a custom control program sized to the company's tools, stage, and risk profile, with a stated path and timeline.

Mapped capabilities

4 capabilities

  • Framework selection

    Recommending the right framework set from the customer's market, data types, and deal blockers.

  • Per-framework path and stages

    Correct stage sequence and duration for the chosen framework, including where a framework differs from its siblings.

  • Control customization

    Tailoring controls to company size, tools, and risk instead of applying a generic checklist.

  • Multi-framework overlap

    Handling shared controls and evidence when more than one framework is in flight.

Illustrative example

Input
We handle PHI and need HIPAA. What are the stages to get compliant and roughly how long does each take?
Expected behavior
The answer gives the HIPAA path — onboarding, platform setup, then BAA collection — and does not insert a three-month observation period or a separate audit window, which belong to other frameworks.

05

Trust Report Surface

The free real-time trust page customers use to show compliance status and security posture to prospects during security review.

Show customers your compliance status and security posture with a real-time trust page. delve.co

Mapped capabilities

3 capabilities

  • Status accuracy

    Published status reflects the real current state, including in-progress versus achieved.

  • Real-time updates

    The page reflects changes in posture rather than a stale snapshot.

  • Disclosure boundaries

    What is shown publicly versus withheld as internal evidence or sensitive configuration.

06

Vendor Review and Audit Workflow

Third-party risk review at scale plus the human-in-the-loop layer: white-glove onboarding, 1:1 Slack support, and end-to-end audit support with the auditor.

We stay with you through the audit, answering questions, providing evidence, and talking to your auditor delve.co

Mapped capabilities

4 capabilities

  • Vendor response analysis

    Auto-analyzing incoming vendor answers, highlighting gaps, and surfacing risk for review.

  • Audit evidence handoff

    Assembling and delivering the evidence an auditor requests, with traceability back to controls.

  • Escalation to humans

    Routing judgment calls to onboarding or Slack support instead of the agent guessing.

  • Task assignment and follow-through

    Assigning remediation work to the right owner and tracking it to closure.

Coverage is mapped from Delve's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Delve test?+

The coverage map is generated from Delve's own public product surface (AI compliance automation platform (SOC 2, HIPAA, GDPR, ISO 27001, PCI-DSS)): 6 scoring areas — Evidence Collection Agents, Questionnaire and Policy Answering, and Code and Infrastructure Scanning, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Delve evals scored?+

Every case generated for Delve — across Evidence Collection Agents and Questionnaire and Policy Answering and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Delve library include?+

The full Delve library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Screenshot and artifact capture and Report drafting under Evidence Collection Agents); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Delve or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Delve areas and set them up in a Corsac workspace, where you can run every test case against Delve or your own agent with your own data.