Eval directory
Evals for Bolt
8 evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Bolt AI products.
About Bolt
Bolt is StackBlitz's AI app builder at bolt.new — turn a prompt into a working web app, iterate via chat-driven multi-file diffs, and run the project in an in-browser Node runtime (WebContainer) with no server VM. Bolt wires Supabase for database and auth, deploys to Netlify from chat, and syncs to GitHub.
Available eval packs for Bolt
8 packs ready to run.
Database And Backend Supabase
Bolt evals — Database & Backend (Supabase) (relift v3 InfraRed)
Deployment And Hosting
Bolt evals — Deployment & Hosting (relift v3 InfraRed)
Github Sync And Code Export
Bolt evals — GitHub Sync & Code Export (relift v3 InfraRed)
Iterative Editing And Diff
Bolt evals — Iterative Editing & Diff (relift v3 InfraRed)
Prompt To App Generation
Bolt evals — Prompt-to-App Generation (relift v3 InfraRed)
Safety Errors And Governance
Bolt evals — Safety, Errors & Governance (relift v3 InfraRed)
Token Credit Economy
Bolt evals — Token / Credit Economy (relift v3 InfraRed)
Webcontainer Runtime
Bolt evals — WebContainer Runtime (relift v3 InfraRed)
Why eval Bolt AI
Bolt's AI features ship behind brand promises about accuracy, safety, and reliability. Buyers and integrators need to know those promises hold up under adversarial prompts, edge-case workflows, and the long tail of real customer inputs — not just the demo path.
The Corsac eval library for Bolt measures four dimensions teams care about most when deploying ai platform agents:
- Adversarial robustness — does the agent resist prompt injection, jailbreaks, and social-engineering attempts?
- Workflow quality— does it complete the task buyers were shown in the demo, on inputs that don't look like the demo?
- Safety gates — does it escalate or refuse when it should, and only then?
- Operator quality — does it preserve analyst trust by surfacing the right context at the right time?
Every eval pack above is hand-authored against Bolt's public product surface and runnable in Corsac with your own data.