
Eval directory
Evals for Bolt
8 evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Bolt AI products.
About Bolt
Bolt is StackBlitz's AI app builder at bolt.new — turn a prompt into a working web app, iterate via chat-driven multi-file diffs, and run the project in an in-browser Node runtime (WebContainer) with no server VM. Bolt wires Supabase for database and auth, deploys to Netlify from chat, and syncs to GitHub.
How complete this published benchmark library is across datasets, metrics, rubrics, use-case maps, and pack context. This is library coverage, not an agent performance score.
Test datasets
8/8 packs
Scoring metrics
0/8 packs
Judge rubrics
8/8 packs
Use-case maps
0/8 packs
Pack context
8/8 packs
Available eval packs for Bolt
8 packs ready to run.
Database And Backend Supabase
Evaluates Bolt's Database & Backend (Supabase) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.
Deployment And Hosting
Evaluates Bolt's Deployment & Hosting across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.
Github Sync And Code Export
Evaluates Bolt's GitHub Sync & Code Export across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.
Iterative Editing And Diff
Evaluates Bolt's Iterative Editing & Diff across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.
Prompt To App Generation
Evaluates Bolt's Prompt-to-App Generation across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.
Safety Errors And Governance
Evaluates Bolt's Safety, Errors & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.
Token Credit Economy
Evaluates Bolt's Token / Credit Economy across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.
Webcontainer Runtime
Evaluates Bolt's WebContainer Runtime across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI App Builder eval coverage.
Why eval Bolt AI
Bolt's AI features ship behind brand promises about accuracy, safety, and reliability. Buyers and integrators need to know those promises hold up under adversarial prompts, edge-case workflows, and the long tail of real customer inputs — not just the demo path.
The Corsac eval library for Bolt measures four dimensions teams care about most when deploying ai platform agents:
- Adversarial robustness — does the agent resist prompt injection, jailbreaks, and social-engineering attempts?
- Workflow quality— does it complete the task buyers were shown in the demo, on inputs that don't look like the demo?
- Safety gates — does it escalate or refuse when it should, and only then?
- Operator quality — does it preserve analyst trust by surfacing the right context at the right time?
Every eval pack above is hand-authored against Bolt's public product surface and runnable in Corsac with your own data.