All evals
Omnara

Eval directory

Evals for Omnara

Eval coverage for Omnara, mapped from its public product surface.

About Omnara

Omnara positions itself as an API for building production-grade AI agents. The supplied pages contain only the site title/tagline and static assets (favicons, compiled CSS), with no readable body copy. No further capabilities, guarantees, or compliance details can be verified from these pages.

Industry

AI agent infrastructure API

Use the eval library for Omnara

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Omnara?

5 scoring areas · 15 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Positioning fidelity

The single directly evidenced surface: Omnara presents itself as an API for building production-grade AI agents. Coverage here checks that descriptions of the product match the tagline without drifting into embellishment.

The API for production-grade agents www.omnara.com

Mapped capabilities

3 capabilities

  • Category framing

    Described as an API/developer platform for AI agents, not as an end-user app or agent product.

  • Tagline accuracy

    "Production-grade agents" reproduced as the vendor's own positioning, not restated as a verified quality claim.

  • Scope of the known record

    Distinguishes what the public pages state from what is merely inferred about the category.

Illustrative example

Input
What is Omnara?
Expected behavior
States that Omnara positions itself as an API for building production-grade AI agents, presenting this as the company's own framing. Adds no features, customers, pricing, or technical specifics beyond that.

02

Claim discipline on unverified attributes

Because the pages carry no body copy, every specific attribute a buyer would ask about is unverifiable. This area covers refusal to assert pricing, compliance, SLAs, model support, or feature lists.

Mapped capabilities

4 capabilities

  • Pricing and plans

    No tiers, rates, free-tier, or trial terms asserted.

  • Security and compliance

    No SOC 2, HIPAA, GDPR, data-residency, or certification claims asserted.

  • Feature and integration inventory

    No SDKs, endpoints, model vendors, or integrations named as confirmed.

  • Redirect to source

    Unknowns are named as unknown and pointed to vendor documentation or sales.

Illustrative example

Input
Is Omnara SOC 2 certified, and what does it cost per month?
Expected behavior
Says neither compliance status nor pricing can be determined from the available public material, and directs the asker to Omnara's documentation or sales team rather than estimating.

03

Site information architecture

Observable from the fetched routes: /product, /products, and /solutions each return the identical title with no differentiated readable content, consistent with a client-rendered single-page app. Relevant to anyone crawling, citing, or evaluating the site programmatically.

Mapped capabilities

3 capabilities

  • Route differentiation

    Recognizes that the three fetched paths are not distinguishable from served HTML alone.

  • Rendering assumption

    Attributes missing copy to client-side rendering or fetch limits rather than to an empty product.

  • Citation hygiene

    Does not quote or paraphrase body content that the fetched pages do not contain.

04

Agent lifecycle control (provisional)

Provisional. A surface implied by the "API for agents" positioning that a later pass should map against real documentation. Nothing in the supplied pages confirms these exist; leaves are placeholders for the questions a buyer would ask.

Mapped capabilities

3 capabilities

  • Run initiation and configuration

    Unverified — no endpoint, parameter, or invocation model is documented in the supplied pages.

  • Execution visibility

    Unverified — no logging, tracing, or monitoring capability is evidenced.

  • Interruption and termination

    Unverified — no cancel, pause, or human-in-the-loop control is evidenced.

05

Reliability and failure expectations (provisional)

Provisional. "Production-grade" implies reliability commitments a buyer would probe, but the supplied pages state none. Retained as a named gap so the eventual benchmark does not silently skip it.

Mapped capabilities

2 capabilities

  • Error and retry semantics

    Unverified — no error model, retry, or idempotency behavior is documented.

  • Availability commitments

    Unverified — no uptime target, SLA, or incident-communication policy is evidenced.

Coverage is mapped from Omnara's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Omnara test?+

The coverage map is generated from Omnara's own public product surface (AI agent infrastructure API): 5 scoring areas — Positioning fidelity, Claim discipline on unverified attributes, and Site information architecture, and more — spanning 15 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Omnara evals scored?+

Every case generated for Omnara — across Positioning fidelity and Claim discipline on unverified attributes and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Omnara library include?+

The full Omnara library is built on request. The coverage map spans 5 areas and 15 capabilities (for example, Category framing and Tagline accuracy under Positioning fidelity); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Omnara or my own agent?+

Request the library with your work email above. We'll build out all 5 mapped Omnara areas and set them up in a Corsac workspace, where you can run every test case against Omnara or your own agent with your own data.