All evals
LS

Eval directory · Security Operations

Evals for Legit Security

Eval coverage for Legit Security, mapped from its public product surface.

About Legit Security

Legit Security is an AI-native ASPM and application security platform that restores visibility and control across engineering tools and workflows so security can keep pace with development. Its VibeGuard product secures AI-generated code at the point of creation, integrating into AI IDEs and code assistants to catch vulnerabilities, secrets, and policy violations before commit. The platform also covers software discovery and inventory, secrets scanning, AI asset visibility, and agentic remediation.

Industry

AI-native application security posture management (ASPM)

Use the eval library for Legit Security

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Legit Security?

6 scoring areas · 22 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

VibeGuard: securing AI-generated code at creation

Real-time scanning and blocking of vulnerabilities, secrets, and policy violations inside AI IDEs and code assistants, before code leaves the developer endpoint.

VibeGuard prevents vulnerabilities, secrets and risk at the developer endpoint www.legitsecurity.com

Mapped capabilities

4 capabilities

  • Pre-commit blocking in the IDE

    Flags or blocks vulnerable AI-generated code at the endpoint before commit.

  • AI code assistant integration

    Operates inside Cursor, GitHub Copilot, and comparable assistants without workflow changes.

  • Real-time policy violation detection

    Surfaces policy breaches as code is generated, not in a later pipeline stage.

  • Low-friction deployment and time-to-value

    Setup measured in minutes with findings visible immediately; no developer slowdown.

Illustrative example

Input
A developer in Cursor accepts an AI-generated Python function that connects to a payments API using a literal key string assigned inline in the source file.
Expected behavior
The secret is flagged at the endpoint before commit, identified as a hardcoded credential with its file and line, and the developer is told to move it to a managed secret rather than being silently allowed through.

02

AI asset visibility

Discovery and inventory of AI components in the engineering estate so security knows where AI is introducing code and capability.

VibeGuard scans AI-generated code for vulnerabilities before it leaves your IDE. www.legitsecurity.com

Mapped capabilities

4 capabilities

  • AI model discovery

    Identifies AI models in use across applications and services.

  • Code assistant discovery

    Detects which assistants developers are running and where.

  • MCP server discovery

    Enumerates MCP servers connected to development workflows.

  • AI-generated code attribution

    Distinguishes AI-generated code from human-written code in the codebase.

03

Software discovery and inventory

Automatic mapping of applications, components, and dependencies across the SDLC, kept current as the environment changes.

ASPM solutions automatically find all your applications, components, and dependencies www.legitsecurity.com

Mapped capabilities

4 capabilities

  • Automatic asset enumeration

    Finds applications, components, and dependencies without manual registration.

  • Continuous change mapping

    Updates inventory as pipelines and repositories change.

  • Blind-spot surfacing

    Highlights forgotten microservices and outdated libraries.

  • Cross-tool correlation

    Connects findings from development to deployment instead of siloing them by tool.

04

Secrets scanning

Detection of credentials and secrets across code and engineering workflows, including at the point of AI code generation.

Automatically detect secrets, issues and policy violations in real-time as developers code. www.legitsecurity.com

Mapped capabilities

3 capabilities

  • Secret detection in generated code

    Catches secrets introduced by AI assistants before commit.

  • Repository and pipeline coverage

    Scans engineering tools and workflows beyond a single repo.

  • Finding quality and noise control

    Presents secrets findings in a form a developer can act on.

05

Risk-based vulnerability prioritization

Deciding what to fix first from actual exposure and exploitability rather than a generic severity score, aligned to CISA BOD 26-04's framing.

Mapped capabilities

4 capabilities

  • Exposure and reachability assessment

    Determines whether a flaw is publicly exposed and reachable.

  • Known-exploited and automation signals

    Accounts for KEV listing and whether exploitation can be automated.

  • Business-context weighting

    Considers proximity to assets the business cares about.

  • Remediation timeline guidance

    Maps a finding to an urgency tier rather than treating every critical alike.

Illustrative example

Input
Two CVSS 9.8 findings: one in a public-facing service and listed in the KEV catalog, one in an internal-only library with no known exploit. Which is fixed first?
Expected behavior
Ranks the exposed KEV-listed finding first and justifies it on exposure and known exploitation, not the shared severity score. Assigns the internal finding a longer timeline instead of treating identical CVSS as identical urgency.

06

Agentic remediation

Automating the fix side of application security, closing the gap between finding a problem and shipping a correction.

Mapped capabilities

3 capabilities

  • Automated fix generation

    Produces a remediation for an identified finding.

  • Developer workflow delivery

    Routes fixes into the tools engineers already use.

  • Remediation gap tracking

    Reports on findings closed versus outstanding.

Coverage is mapped from Legit Security's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Legit Security test?+

The coverage map is generated from Legit Security's own public product surface (AI-native application security posture management (ASPM)): 6 scoring areas — VibeGuard: securing AI-generated code at creation, AI asset visibility, and Software discovery and inventory, and more — spanning 22 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Legit Security evals scored?+

Every case generated for Legit Security — across VibeGuard: securing AI-generated code at creation and AI asset visibility and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Legit Security library include?+

The full Legit Security library is built on request. The coverage map spans 6 areas and 22 capabilities (for example, Pre-commit blocking in the IDE and AI code assistant integration under VibeGuard: securing AI-generated code at creation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Legit Security or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Legit Security areas and set them up in a Corsac workspace, where you can run every test case against Legit Security or your own agent with your own data.