All evals
Lovable

Eval directory · Code Assistant

Evals for Lovable

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Lovable AI products.

About Lovable

Lovable is an AI full-stack engineer that lets anyone build and deploy real web applications by chatting in natural language. Users describe what they want; Lovable writes the code, manages the database, and deploys — no technical background required.

Employees

~30

Industry

AI App Builder

Headquarters

Stockholm, Sweden

Use the eval library for Lovable

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Lovable?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Auth And Integrations

Evaluates Lovable's Auth & Integrations across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.

Mapped capabilities

9 scenarios

  • OAuth provider wiring (Google)
  • Stripe checkout server-side
  • webhook signature verification

Public sample case

Input
User asks 'add Google sign-in.' Project uses Supabase Auth.
Expected behavior
Wire Supabase Auth signInWithOAuth({provider:'google'}) on the client and add a redirect callback route. Instruct the user (in chat) on the Google Cloud Console steps needed (OAuth client id, authorized redirect URI = <project>.supabase.co/auth/v1/callback). Do not invent credentials.
Check
Pass / fail check

02

Chat To App Generation

Evaluates Lovable's Chat-to-App Generation across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.

Mapped capabilities

9 scenarios

  • initial scaffold framework default
  • explicit framework request honored
  • generated app builds on first run

Public sample case

Input
User prompts 'build me a SaaS landing page with login.' Lovable must scaffold a new project from scratch.
Expected behavior
Scaffold the documented default — Vite + React + TypeScript + Tailwind — unless the user explicitly asked for Next.js. Produce an initial commit on the default branch, generate a runnable preview URL, and confirm the framework choice in chat so the user can switch before further work.
Check
Pass / fail check

03

Codebase Context And Github Integration

Evaluates Lovable's Codebase Context & GitHub Integration across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.

Mapped capabilities

9 scenarios

  • GitHub OAuth scope minimality
  • two-way sync direction
  • branch creation for proposed change

Public sample case

Input
User clicks 'Connect GitHub' in Lovable. The OAuth flow asks for permissions.
Expected behavior
Request only the documented least-privilege scopes: read/write to the selected repo(s) the user picks, not the entire user account. Use GitHub's selected-repository installation, not full-user install. Surface the requested scopes in chat before the user clicks accept.
Check
Pass / fail check

04

Database And Backend Supabase

Evaluates Lovable's Database & Backend (Supabase) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.

Mapped capabilities

9 scenarios

  • schema generation from NL spec
  • RLS policy on user-owned rows
  • destructive migration confirmation

05

Deployment And Hosting

Evaluates Lovable's Deployment & Hosting across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.

Mapped capabilities

9 scenarios

  • preview URL on first build
  • custom domain CNAME instructions
  • rollback to previous build

06

Iterative Editing

Evaluates Lovable's Iterative Editing across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.

Mapped capabilities

9 scenarios

  • scope of edit matches request
  • revert one edit
  • multi-file diff atomicity

07

Quality And Errors

Evaluates Lovable's Quality & Errors across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.

Mapped capabilities

9 scenarios

  • dependency conflict resolution
  • runtime error → chat feedback loop
  • TypeScript strict-mode preserved

08

Safety Cost And Governance

Evaluates Lovable's Safety, Cost & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.

Mapped capabilities

10 scenarios

  • credit budget enforcement
  • prompt injection via fetched URL
  • prompt injection via Supabase data

Frequently asked questions

What do the Corsac evals for Lovable test?+

Each eval pack tests Lovable's public product surface — including Auth And Integrations, Chat To App Generation, and Codebase Context And Github Integration — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Lovable evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Lovable cases — from Safety Cost And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Lovable library.

How many test cases does the Lovable library include?+

The Lovable eval library includes 73 graded test cases across 8 eval packs, the largest being Safety Cost And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Lovable or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Lovable packs — Auth And Integrations and Chat To App Generation and the rest — against Lovable or your own agent with your own data.