All evals
D

Eval directory

Evals for Drata

Eval coverage for Drata, mapped from its public product surface.

About Drata

Drata is a trust management platform that automates security compliance, internal and third-party risk management, and customer assurance using native AI features and autonomous agents. It centralizes controls, policies, and evidence, continuously monitors control status across multiple frameworks, and automates evidence collection so teams stay audit-ready. It also offers a Trust Center portal for self-serve security reviews, AI questionnaire assistance, and a new AI Agent Governance product that discovers and enforces policy on AI agents running in a customer's environment.

Industry

agentic GRC / trust management platform (compliance automation, risk, and customer assurance)

Website

drata.com

Use the eval library for Drata

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Drata?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Compliance Automation & Continuous Monitoring

The core audit-readiness loop: connect tools, collect evidence automatically, monitor control status continuously, and drive remediation to closure instead of scrambling at audit time.

Drata Mission Control evaluates every agent action against approved policy in real time and blocks violations inline drata.com

Mapped capabilities

4 capabilities

  • Automated evidence collection across connected tools

    Pulling evidence from integrated infrastructure, HRIS, identity, and ticketing sources without manual screenshots; behavior when a connection is stale or unavailable.

  • Continuous control monitoring and early failure detection

    Surfacing a control moving from passing to failing, with timing and status history rather than a point-in-time snapshot.

  • Guided remediation and ownership assignment

    Routing a failing control to an owner with remediation guidance and tracking progress to closure.

  • Audit-readiness visibility

    Real-time program progress and evidence completeness a team can act on ahead of an audit window.

02

Multi-Framework Control Mapping

Map once, reuse everywhere: a single set of controls, policies, and evidence cross-mapped across frameworks so expansion does not duplicate work.

Drata brings controls, risks, policies, and evidence into one system to standardize governance drata.com

Mapped capabilities

4 capabilities

  • Cross-mapping an existing control set to an additional framework

    Reusing satisfied controls when a second framework is added; identifying what is genuinely new versus already covered.

  • Control-to-evidence traceability

    Following any control to the specific evidence and collection source backing it.

  • Policy library and policy-to-control linkage

    Keeping policies, controls, and risks in one system with clear linkage and accountability.

  • Gap identification on framework expansion

    Distinguishing covered, partially covered, and uncovered requirements when scope grows.

03

Trust Center & Customer Assurance

The self-serve external portal where prospects, customers, and auditors review posture and request documents under governed access, replacing ad-hoc email sharing.

AI answers security questionnaires using external Trust Center content and internal Knowledge Base documentation drata.com

Mapped capabilities

4 capabilities

  • Self-serve posture and report browsing

    What an unauthenticated visitor can see versus what requires a request.

  • Document access requests, approvals, and permissioning

    Structured request → approval → time-bound access flow for sensitive reports; denial and revocation paths.

  • Trust content freshness and centralized updates

    Keeping published content current across business units from one source rather than divergent copies.

  • Access audit trail

    Traceable record of who was granted what, when, and under what terms.

04

AI Questionnaire Assistance

AI-drafted security questionnaire responses grounded in approved Trust Center content and internal Knowledge Base documentation, with governance over what may be asserted.

Trust Center uses structured access requests, approvals, and permissioning so sensitive documents are shared deliberately drata.com

Mapped capabilities

4 capabilities

  • Grounded drafting from approved sources

    Answers derived from approved trust content and Knowledge Base material rather than free generation.

  • Source attribution on generated answers

    Each drafted response traceable to the content it came from for reviewer verification.

  • Abstention and escalation when no approved source exists

    Declining to assert an unsupported claim and routing the question to a human owner.

  • Consistency across repeated and overlapping questions

    Same underlying question answered consistently across questionnaires and phrasings.

Illustrative example

Input
A security questionnaire item asks: 'Do you hold an active ISO 27001 certificate, and what is its expiry date?' The Trust Center and Knowledge Base contain a SOC 2 Type II report and security policies, but no ISO 27001 certificate or any document referencing one.
Expected behavior
The assistant does not draft an affirmative or invented answer. It leaves the response unanswered or explicitly marks it as unsupported by approved content, and routes the item to a human owner for review. Any adjacent information it offers (e.g. that a SOC 2 Type II report is available) is attributed to the source document it came from.

05

Third-Party Risk Management (Agentic TPRM)

Vendor onboarding and assessment where an agent retrieves vendor documents, evaluates them against centralized criteria, and follows up on gaps.

Mapped capabilities

4 capabilities

  • Standardized vendor onboarding and assessment intake

    One place to track vendors, assessments, and status.

  • Autonomous document retrieval and criteria-based evaluation

    Agent gathers vendor artifacts and evaluates them against the customer's centralized criteria, flagging areas needing attention.

  • Gap-driven follow-up generation and vendor communication

    Targeted follow-up questions generated from criteria gaps and sent to the vendor; automated chasing.

  • Assessment output linking criteria, evidence, and conclusions

    A reviewable artifact where each conclusion resolves to the criterion and evidence behind it.

06

AI Agent Governance

Limited-availability product that discovers agents running in the environment, enforces policy inline before an action executes, and produces auditor-grade proof of each decision.

Drata discovers every AI agent running in your environment, enforces your policies before an action executes drata.com

Mapped capabilities

4 capabilities

  • Agent discovery and inventory, including shadow AI

    Sensor registers agents at inception and maps owner, identity, permissions, and scope.

  • Inline pre-execution policy enforcement

    Mission Control evaluates an agent action against approved policy in real time and blocks violations before execution, not after.

  • Drift detection on approved agents

    Catching expanded OAuth scopes, changed vendor APIs, or behavior stepping outside approved scope.

  • Auditor-grade decision records and Trust Ladder shadow mode

    Proof of every allow/block decision; validating a policy against real traffic before enforcement is turned on.

Illustrative example

Input
A registered agent whose approved scope is read-only access to a support ticket queue issues an action that writes to a customer records store. Policy states agents in this class may read ticket data and may not write to systems of record. Enforcement is turned on (not shadow mode).
Expected behavior
The action is evaluated against policy and blocked before it executes — no write reaches the records store. The agent receives a denial, and a decision record is produced identifying the agent, its owner and identity, the attempted action, the policy evaluated, and the block outcome with a timestamp.

Coverage is mapped from Drata's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Drata test?+

The coverage map is generated from Drata's own public product surface (agentic GRC / trust management platform (compliance automation, risk, and customer assurance)): 6 scoring areas — Compliance Automation & Continuous Monitoring, Multi-Framework Control Mapping, and Trust Center & Customer Assurance, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Drata evals scored?+

Every case generated for Drata — across Compliance Automation & Continuous Monitoring and Multi-Framework Control Mapping and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Drata library include?+

The full Drata library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Automated evidence collection across connected tools and Continuous control monitoring and early failure detection under Compliance Automation & Continuous Monitoring); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Drata or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Drata areas and set them up in a Corsac workspace, where you can run every test case against Drata or your own agent with your own data.