
Deployment And Hosting
Lovable · Lovable
AI App Builder — Lovable
Evaluates Lovable's Deployment & Hosting across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.
About Lovable
Lovable is an AI full-stack engineer that lets anyone build and deploy real web applications by chatting in natural language. Users describe what they want; Lovable writes the code, manages the database, and deploys — no technical background required.
Sample tests· showing 3 of 9
| # | Input | Expected behavior | Check |
|---|---|---|---|
| 01 | Scaffold finished. User wants the preview URL to share. | Deploy the build to <slug>.lovable.app, surface the URL in chat and the project dashboard. The URL should serve the just-built artifact (not a stale or empty page) and be reachable from any browser without login. | Pass / FailAi Platformhigh |
| 02 | Latest build introduced a runtime regression. User clicks 'rollback.' | Restore the previous deployed build artifact to the preview URL within the documented rollback path, keep the source tree untouched (or revert + commit on a branch if source rollback is desired), and confirm the artifact serving the URL. | Pass / FailAi Platformcritical |
| 03 | User wants to point app.acme.com at their Lovable app. | Show the documented CNAME target in the UI / chat, request DNS verification, and wait for propagation before issuing HTTPS. Confirm HTTPS is automatic per docs or surface the manual step. Do not claim 'done' before verification completes. | Pass / FailAi Platformhigh |
How this eval is graded
Grade against expected.ideal_behavior and expected.rubric. Per-criterion pass requires mean >= 4.0 and no criterion below 3.
Rubric criteria
- Lovable
- Ai Platform
- Deployment And Hosting
Recommended for
Works with
Related evals
Browserbase
Evaluates Browserbase's Captcha Handling across scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser infrastructure eval coverage.
View Code AssistantBrowserbase
Evaluates Browserbase's Concurrency & Rate Limits across scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser infrastructure eval coverage.
View Code AssistantBrowserbase
Evaluates Browserbase's Live Debugging & Session Inspector across scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Browser infrastructure eval coverage.
ViewFrequently asked questions
What does the Deployment And Hosting eval for Lovable Lovable test?+
Evaluates Lovable's Deployment & Hosting across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.
How is the Deployment And Hosting eval scored?+
The judge rubric: Grade against expected.ideal_behavior and expected.rubric. Per-criterion pass requires mean >= 4.0 and no criterion below 3.
How many test cases does this eval pack include?+
The Deployment And Hosting pack for Lovable Lovable contains 9 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.
How do I run this eval?+
Sign up for Corsac, connect your model or agent endpoint, and run the Deployment And Hosting pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.
Run this eval in your workspace
Connect your data, configure thresholds, and review results with your team.