All evals
S

Eval directory

Evals for Softgen

Eval coverage for Softgen, mapped from its public product surface.

About Softgen

Softgen is an AI web app builder that turns a natural-language description into a full-stack web application. It offers a choice of multiple AI models, managed Supabase for backend/data, and code ownership for the user. New users can start on a free trial with credits and no credit card.

Industry

AI web app builder

Website

softgen.ai

Use the eval library for Softgen

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Softgen?

6 scoring areas · 21 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Natural-Language App Generation

The core promise: a user describes what they want and Softgen produces a working full-stack web app, then refines it across follow-up turns.

Trial credits drop instantly — start building your first full-stack app in minutes. softgen.ai

Mapped capabilities

4 capabilities

  • Initial build from a plain-language description

    Turning an unstructured 'what I want to build' prompt into a scoped full-stack app rather than a partial page.

  • Iterative refinement across sessions

    Applying follow-up change requests to an existing app without discarding prior work.

  • Ambiguous or underspecified requests

    Handling prompts that omit key details by asking or stating assumptions instead of silently guessing.

  • Scope boundaries of a web app build

    Distinguishing what the builder produces (full-stack web apps) from requests outside that surface.

02

AI Model Choice

Softgen exposes 12+ AI models (including Claude, GPT-5.5, and Gemini families) and lets builders choose among them for a given project.

AI Models for App Building: Claude, GPT-5.5, Gemini softgen.ai

Mapped capabilities

3 capabilities

  • Accurate roster of available models

    Naming only models Softgen actually offers, without inventing vendors or versions.

  • Model selection guidance for a task

    Recommending among offered models for a stated build without fabricating benchmark claims.

  • Switching models on an existing project

    Explaining whether and how a builder changes model mid-project.

Illustrative example

Input
I'm building a data-heavy admin dashboard. Which AI model should I pick in Softgen, and can I switch to a different one later?
Expected behavior
Recommends from the models Softgen actually offers, such as the Claude, GPT-5.5, or Gemini options, and confirms the choice spans 12+ models the builder can select. Gives reasoning without citing benchmark numbers.

03

Managed Supabase Backend & Data

Softgen provisions and manages Supabase as the backend and data layer for generated apps.

Managed Supabase softgen.ai

Mapped capabilities

3 capabilities

  • Backend and schema provisioning

    Standing up data storage for an app described in natural language.

  • Auth and user accounts in generated apps

    Wiring sign-in and per-user data access as part of the managed backend.

  • Explaining what 'managed' covers

    Being clear about which backend responsibilities Softgen handles versus the builder.

04

Code Ownership & Portability

Softgen's 'your code, forever' commitment — builders retain the generated application code.

Your code, forever softgen.ai

Mapped capabilities

3 capabilities

  • Accurate ownership claims

    Stating code ownership as advertised without overstating rights not evidenced on the site.

  • Getting the code out

    Answering how a builder accesses or takes the generated codebase with them.

  • Continuity if the builder leaves Softgen

    Addressing what a user keeps when they stop using the product.

05

Trial, Credits & Plans

Onboarding runs on a free trial with starting credits and no credit card; ongoing usage is credit-based across paid plans.

Start free with trial credits — no card required. softgen.ai

Mapped capabilities

4 capabilities

  • Free trial terms

    No credit card required and trial credits granted at signup.

  • How credits are consumed

    Explaining that builds and iterations draw down credits.

  • Plan and pricing questions

    Routing to the pricing page rather than inventing tier names or dollar figures.

  • Running out of credits

    What happens to an in-progress project when the trial balance is exhausted.

Illustrative example

Input
Is Softgen free to try? Do I need to enter a credit card, and how many credits do I get to start?
Expected behavior
Confirms a free trial with no credit card required and trial credits granted at signup, citing the 75 starting credits. Directs the user to the pricing page for paid plans instead of quoting tier names or dollar amounts.

06

Policy & Comparative Claims

Softgen publishes legal terms, an abuse policy, a changelog, and head-to-head comparison pages against Lovable, Bolt.new, v0, Replit, and Base44.

Mapped capabilities

4 capabilities

  • Abuse and acceptable-use boundaries

    Declining build requests that fall outside the published abuse policy.

  • Competitor comparisons

    Comparing against the five named products fairly, without claims absent from the compare pages.

  • Changelog and shipped-feature accuracy

    Describing capabilities as shipped only when the changelog or product pages support it.

  • Testimonial and outcome claims

    Treating community results as individual anecdotes rather than guaranteed outcomes.

Coverage is mapped from Softgen's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Softgen test?+

The coverage map is generated from Softgen's own public product surface (AI web app builder): 6 scoring areas — Natural-Language App Generation, AI Model Choice, and Managed Supabase Backend & Data, and more — spanning 21 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Softgen evals scored?+

Every case generated for Softgen — across Natural-Language App Generation and AI Model Choice and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Softgen library include?+

The full Softgen library is built on request. The coverage map spans 6 areas and 21 capabilities (for example, Initial build from a plain-language description and Iterative refinement across sessions under Natural-Language App Generation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Softgen or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Softgen areas and set them up in a Corsac workspace, where you can run every test case against Softgen or your own agent with your own data.