All evals
WOMBO

Eval directory

Evals for WOMBO

Eval coverage for WOMBO, mapped from its public product surface.

About WOMBO

Dream by WOMBO is a consumer-facing AI art and video generator. The provided pages describe it as creating images and videos from text prompts. The captured pages contain only title, manifest, icon, and stylesheet data, so no further capability, pricing, or compliance details are available.

Industry

consumer AI image & video generation app

Website

dream.ai

Use the eval library for WOMBO

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for WOMBO?

5 scoring areas · 20 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Text-to-Image Generation

The core advertised capability: turning a written prompt into a still image. Covers whether a prompt reliably produces a returned image asset whose subject matches what was asked for.

create stunning images and videos from text prompts dream.ai

Mapped capabilities

4 capabilities

  • Prompt-to-image fulfillment

    A submitted text prompt yields at least one retrievable, valid image asset.

  • Subject fidelity

    The generated image depicts the primary subject named in the prompt.

  • Multi-element prompts

    Prompts naming more than one subject or attribute are handled without dropping a stated element.

  • Empty or degenerate prompts

    Blank, whitespace-only, or extremely short prompts produce a clear outcome rather than a silent failure.

Illustrative example

Input
Submit the prompt "a red fox sitting in falling snow" and request an image.
Expected behavior
The generation completes and returns at least one retrievable image asset. The image depicts a fox as its primary subject in a snowy setting, rather than an unrelated scene, a blank canvas, or a text-only response.

02

Text-to-Video Generation

The second advertised output mode. Covers whether a prompt produces a playable video asset and whether motion output is distinguishable from a still image.

Mapped capabilities

4 capabilities

  • Prompt-to-video fulfillment

    A submitted prompt yields a retrievable video asset in a playable container.

  • Output mode selection

    Requesting video versus image returns the corresponding media type.

  • Motion presence

    Video output contains more than one distinct frame.

  • Prompt-to-video subject fidelity

    The clip depicts the subject named in the prompt.

03

Prompt Interpretation and Creative Control

How the product reads the user's wording. Covers the mapping from natural-language phrasing to the resulting artwork, which is the main lever a consumer has over output.

Mapped capabilities

4 capabilities

  • Style descriptors

    Style words in the prompt are reflected in the rendered result.

  • Prompt length tolerance

    Both terse and long prompts are accepted and produce output.

  • Non-English and mixed-script prompts

    Prompts in other scripts are accepted without input rejection.

  • Repeatability of a fixed prompt

    Resubmitting the same prompt returns an on-topic result each time.

04

Installable Web App Shell

The manifest declares an installable, standalone PWA. Covers whether the install metadata is well-formed and self-consistent, which determines whether the product can be added to a home screen at all.

Mapped capabilities

4 capabilities

  • Manifest validity

    site.webmanifest parses as JSON and carries name, short_name, start_url, scope, and display.

  • Icon availability

    Every declared icon source resolves and matches its stated size and MIME type.

  • Standalone launch

    start_url loads the app within the declared scope in standalone display mode.

  • Entry-point equivalence

    The bare host and the trailing-slash root serve the same application.

Illustrative example

Input
Fetch https://dream.ai/site.webmanifest, then fetch every icon source it declares.
Expected behavior
The manifest parses as JSON and declares the product name, a standalone display mode, and a start_url inside its own scope. Each declared icon resolves successfully at the exact pixel dimensions and MIME type the manifest states.

05

Visual Shell Consistency

The captured stylesheet and manifest pin a specific dark, typographic identity. Covers whether the delivered shell matches the declared brand tokens across surfaces a user first sees.

Mapped capabilities

4 capabilities

  • Theme and background color

    Rendered chrome matches the declared theme and background colors.

  • Typography

    The declared primary typeface loads, with the stated fallback stack applied when it does not.

  • Favicon set

    All referenced favicon and touch-icon assets resolve as valid images.

  • Document title

    The served page title matches the declared product name.

Coverage is mapped from WOMBO's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for WOMBO test?+

The coverage map is generated from WOMBO's own public product surface (consumer AI image & video generation app): 5 scoring areas — Text-to-Image Generation, Text-to-Video Generation, and Prompt Interpretation and Creative Control, and more — spanning 20 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the WOMBO evals scored?+

Every case generated for WOMBO — across Text-to-Image Generation and Text-to-Video Generation and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the WOMBO library include?+

The full WOMBO library is built on request. The coverage map spans 5 areas and 20 capabilities (for example, Prompt-to-image fulfillment and Subject fidelity under Text-to-Image Generation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against WOMBO or my own agent?+

Request the library with your work email above. We'll build out all 5 mapped WOMBO areas and set them up in a Corsac workspace, where you can run every test case against WOMBO or your own agent with your own data.