All evals
Agentastic

Eval directory

Evals for Agentastic

Eval coverage for Agentastic, mapped from its public product surface.

About Agentastic

Agentastic is a company whose site is titled "Agentastic - Building the Collective Intelligence." The retrieved pages contain only the site title/tagline plus static assets (favicons, JavaScript bundles, CSS, and font files), with no readable product content. No product category, capabilities, or company facts can be determined from these pages.

Use the eval library for Agentastic

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Agentastic?

4 scoring areas · 12 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Public identity and tagline fidelity

The only human-readable content retrieved is the site title and tagline. This area covers whether an assistant reproduces that identity exactly and attributes it to the right domain, without embellishment.

Mapped capabilities

3 capabilities

  • Exact tagline reproduction

    Returns "Building the Collective Intelligence" verbatim rather than a paraphrase or expansion.

  • Name and domain binding

    Associates the name Agentastic with agentastic.ai and does not conflate it with similarly named entities.

  • Title-tag vs. product claim separation

    Presents the tagline as marketing copy, not as a description of a shipped capability.

Illustrative example

Input
What is the tagline on agentastic.ai, exactly as written?
Expected behavior
Returns the tagline "Building the Collective Intelligence" exactly as it appears in the page title, attributed to agentastic.ai, without paraphrase, expansion, or added interpretation of what it means.

02

Insufficient-evidence handling

The dominant characteristic of this surface is absence of information. This area covers whether an assistant recognizes and states that the retrieved pages do not establish a product category, and refuses to fill the gap.

Mapped capabilities

3 capabilities

  • Explicit unknown declaration

    States that product category and capabilities cannot be determined from the retrieved pages.

  • No inferred product category

    Does not infer an AI-agent product, platform, or feature set from the name or tagline alone.

  • Next-step routing

    Points to what additional evidence would be required rather than speculating.

Illustrative example

Input
Based only on the pages retrieved from agentastic.ai, what product does Agentastic sell and who is it for?
Expected behavior
States that the retrieved pages contain only the site title/tagline and static assets, so the product category and audience cannot be determined, and identifies what further evidence would answer the question. Offers no guessed category.

03

Brand asset surface

Three icon assets are exposed at conventional paths. This area covers accurate description of what brand assets are published and at which URLs, without asserting anything about their visual content.

Mapped capabilities

2 capabilities

  • Icon path inventory

    Identifies /favicon.ico, /favicon.png, and /apple-touch-icon.png as the published icon set.

  • Binary-content restraint

    Does not describe logo imagery, colors, or shapes that were not decoded from the byte payloads.

04

Web delivery and build stack

Asset URLs and bundle contents disclose the frontend stack and deployment scheme. This area covers correct, evidence-bounded characterization of that delivery surface.

Mapped capabilities

4 capabilities

  • Framework identification

    Recognizes Next.js from /_next/static paths and Turbopack from the chunk runtime, with no version claim.

  • Deploy fingerprint interpretation

    Treats the shared dpl_ query parameter as a cache-busting deployment id, not as user-specific or sensitive data.

  • Typography inventory

    Identifies Fira Code woff2 faces with weight range 300-700 and font-display: swap from the CSS bundle.

  • Rendering-mode restraint

    Avoids concluding server- vs. client-rendering behavior from bundle strings such as the client-side-rendering bailout error.

Coverage is mapped from Agentastic's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Agentastic test?+

The coverage map is generated from Agentastic's own public product surface: 4 scoring areas — Public identity and tagline fidelity, Insufficient-evidence handling, and Brand asset surface, and more — spanning 12 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Agentastic evals scored?+

Every case generated for Agentastic — across Public identity and tagline fidelity and Insufficient-evidence handling and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Agentastic library include?+

The full Agentastic library is built on request. The coverage map spans 4 areas and 12 capabilities (for example, Exact tagline reproduction and Name and domain binding under Public identity and tagline fidelity); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Agentastic or my own agent?+

Request the library with your work email above. We'll build out all 4 mapped Agentastic areas and set them up in a Corsac workspace, where you can run every test case against Agentastic or your own agent with your own data.