All evals
Corvic AI

Eval directory

Evals for Corvic AI

Eval coverage for Corvic AI, mapped from its public product surface.

About Corvic AI

Corvic AI is an enterprise platform for composing agents and multi-agent workflows over complex company data. It wraps frontier models with grounded playbooks, persistent memory, a secure vault, and lineage/guardrails so teams can build intelligent applications without maintaining custom pipelines. It is sold in Developer, Premium, and Enterprise tiers with usage-based overages.

Industry

enterprise AI data/agent orchestration platform

Use the eval library for Corvic AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Corvic AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agent & Workflow Composition

Building agents, curated playbooks, and scheduled multi-agent workflows over company data, including model selection across the supported frontier models and MCP tool integration.

workflow-level orchestration substantially outperforms direct single-pass LLM access on detailed engineering extraction tasks www.corvic.ai

Mapped capabilities

4 capabilities

  • Playbook-grounded agent configuration

    Composing an agent from a grounded playbook and its attached tools rather than ad hoc prompting.

  • Multi-agent workflow orchestration

    Sequencing and scheduling multi-step workflows, including handoff of intermediate state between steps.

  • Model selection and comparison

    Choosing among available frontier models (e.g. Claude Sonnet 4.6, GPT 5.6 Terra/Luna, GLM 5.2, DeepSeek V4 Flash, Grok 4.5) and surfacing the comparison view.

  • MCP and native tool invocation

    Invoking MCP integrations and native tools from within a room or agent session.

02

Grounded Answers & Lineage

Reasoning over combined modalities — documents, structured data, knowledge graph relationships, and live web context — with traceable answers rather than unsupported generation.

Unlimited context memory persistent enterprise knowledge www.corvic.ai

Mapped capabilities

4 capabilities

  • Multi-modal retrieval in one reasoning layer

    Answering from PDFs, structured databases, and graph relationships together instead of forcing a single modality.

  • Citation and lineage attribution

    Attaching source lineage to each claim so a reviewer can trace an answer back to its underlying data.

  • Live web context blending

    Incorporating web search results alongside internal data and distinguishing the two in the answer.

  • Abstention on ungrounded questions

    Declining or flagging when connected data does not support the requested claim.

Illustrative example

Input
In your P&ID extraction benchmark, which Corvic configuration was tested and what average client format compliance did it score?
Expected behavior
Reports Opus 4.8 accessed via a deployed workflow at 91.0% average client format compliance, and attributes the figure to the published June 2026 P&ID-to-XML benchmark report rather than asserting it unsourced.

03

Persistent Memory & Reproducibility

Durable state across sessions so results are auditable and repeatable, and shared context is reused rather than re-sent on every call.

Usage beyond your plan's included limits is billed automatically at the rates below. www.corvic.ai

Mapped capabilities

4 capabilities

  • Cross-session knowledge persistence

    Retaining enterprise knowledge so users do not re-upload the same files before every session.

  • Long-context retention beyond window limits

    Maintaining answer quality on material larger than a single model context window.

  • Reproducible answers and audit trail

    Returning consistent results for the same prompt and input, with a retrievable record of the run.

  • Token reuse across pipeline steps

    Persisting intermediate state so shared context is not re-injected at each step of a multi-step run.

04

Security, Governance & Guardrails

Controls that make enterprise deployment viable: secure vault, zero data retention, access and identity, and policy guardrails against shadow AI and data leakage.

The AI operating system for your most complex enterprise data. www.corvic.ai

Mapped capabilities

4 capabilities

  • Secure vault and ZDR handling

    Keeping customer data private and honoring zero-data-retention expectations for connected sources.

  • SSO and access control

    Enterprise identity, seat-level access, and tier-gated administrative capabilities.

  • Guardrail enforcement on agent actions

    Blocking or escalating agent behavior that violates configured policy.

  • Compliance-oriented evidence

    Producing the run records and lineage that audit-style reviews require.

05

Plans, Entitlements & Usage Metering

Tier entitlements and usage-based overage accounting across Developer ($20/mo), Premium ($200/mo), and Enterprise, plus trial handling.

Mapped capabilities

4 capabilities

  • Tier entitlement limits

    Seats, AI tokens, and web search queries included at each published tier.

  • Overage calculation

    Billing usage beyond included limits at $20/extra seat/month, $5/1M tokens, and $20/1,000 web search queries.

  • Trial and upgrade flow

    7-day free trial entry, start-free-trial paths, and contact-sales routing for Enterprise.

  • Usage visibility

    Reporting consumption against included limits before overage is incurred.

Illustrative example

Input
I'm on the Premium tier. This month I used 63M AI tokens and 7,000 web search queries with my included 10 seats. What will I be billed?
Expected behavior
Applies the published Premium base of $200/month, then bills 13M overage tokens at $5 per 1M ($65) and 2,000 overage web search queries at $20 per 1,000 ($40), for a $305 total, with no seat overage.

06

Structured Document Extraction

Converting complex source documents into a client-specified structured schema through a deployed workflow, as evidenced by the published P&ID-to-XML benchmark.

Corvic ranked first on both sheets, averaging 91.0% client format compliance www.corvic.ai

Mapped capabilities

4 capabilities

  • Client schema format compliance

    Emitting output that conforms to a supplied reference XML template.

  • Entity extraction fidelity

    Capturing equipment, instruments, and piping from annotated engineering drawings.

  • Workflow-orchestrated vs single-pass extraction

    Routing extraction through a deployed workflow rather than a direct single-pass model call.

  • Extraction run reporting

    Reporting processing time and per-category results for a completed extraction run.

Coverage is mapped from Corvic AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Corvic AI test?+

The coverage map is generated from Corvic AI's own public product surface (enterprise AI data/agent orchestration platform): 6 scoring areas — Agent & Workflow Composition, Grounded Answers & Lineage, and Persistent Memory & Reproducibility, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Corvic AI evals scored?+

Every case generated for Corvic AI — across Agent & Workflow Composition and Grounded Answers & Lineage and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Corvic AI library include?+

The full Corvic AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Playbook-grounded agent configuration and Multi-agent workflow orchestration under Agent & Workflow Composition); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Corvic AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Corvic AI areas and set them up in a Corsac workspace, where you can run every test case against Corvic AI or your own agent with your own data.