All evals
Vercel

Eval directory · Code Assistant

Evals for Vercel

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Vercel AI products.

About Vercel

Vercel is the frontend cloud platform powering Next.js and modern web deployments. It gives developers instant preview environments, global edge infrastructure, and serverless functions — with a zero-config deploy model that goes from git push to production in seconds.

Employees

~500

Industry

Developer Platform

Headquarters

San Francisco, CA

Website

vercel.com

Use the eval library for Vercel

All 65 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Vercel?

8 areas · 65 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Build Performance

Evaluates Vercel's Build Performance across 8 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Frontend Cloud & Edge Infrastructure eval coverage.

Mapped capabilities

8 scenarios

  • turbo filter
  • file tracing
  • install command

Public sample case

Input
turbo run build runs 40 packages; PR touches apps/web only.
Expected behavior
Agent uses turbo --filter=...[origin/main] or similar in vercel build command, documents affected graph.
Check
Pass / fail check

02

Custom Domains Tls

Evaluates Vercel's Custom Domains & TLS across 7 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Frontend Cloud & Edge Infrastructure eval coverage.

Mapped capabilities

7 scenarios

  • apex www redirect
  • cert renewal
  • multi domain

Public sample case

Input
Apex must 308 to www.example.com; both added in Vercel Domains.
Expected behavior
Agent configures redirect in Domains UI or vercel.json redirects, verifies apex ALIAS/ A and www CNAME per docs.
Check
Pass / fail check

03

Deployment Pipeline

Evaluates Vercel's Deployment Pipeline across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Frontend Cloud & Edge Infrastructure eval coverage.

Mapped capabilities

9 scenarios

  • promote to production
  • rollback
  • CLI prod deploy

Public sample case

Input
Team merged PR to main; Vercel Git integration shows deployment dpl_abc in BUILDING for 40m. Production domain still serves previous deployment. Project has disabled auto-promotion; no one called promote API.
Expected behavior
Agent verifies latest successful build for main, uses POST /v13/deployments/{id}/promote or Dashboard Promote to Production, and documents that auto-promotion was off.
Check
Pass / fail check

04

Edge Middleware

Evaluates Vercel's Edge Middleware across 8 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Frontend Cloud & Edge Infrastructure eval coverage.

Mapped capabilities

8 scenarios

  • geo routing
  • Edge Config AB
  • auth gate

05

Preview Environments

Evaluates Vercel's Preview Environments across 7 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Frontend Cloud & Edge Infrastructure eval coverage.

Mapped capabilities

7 scenarios

  • deployment protection
  • commit preview URL
  • custom environment

06

Serverless Functions

Evaluates Vercel's Serverless Functions across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Frontend Cloud & Edge Infrastructure eval coverage.

Mapped capabilities

9 scenarios

  • maxDuration
  • payload limit
  • cold start archive

07

Team Access Observability

Evaluates Vercel's Team Access & Observability across 8 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Frontend Cloud & Edge Infrastructure eval coverage.

Mapped capabilities

8 scenarios

  • SAML SSO
  • offboarding
  • audit log

08

V0 Ai Code Generation

Evaluates Vercel's v0 AI Code Generation across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Frontend Cloud & Edge Infrastructure eval coverage.

Mapped capabilities

9 scenarios

  • XSS in generated UI
  • secret in client
  • SSRF fetch

Frequently asked questions

What do the Corsac evals for Vercel test?+

Each eval pack tests Vercel's public product surface — including Build Performance, Custom Domains Tls, and Deployment Pipeline — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Vercel evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 65 Vercel cases — from Deployment Pipeline (9 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Vercel library.

How many test cases does the Vercel library include?+

The Vercel eval library includes 65 graded test cases across 8 eval packs, the largest being Deployment Pipeline with 9 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Vercel or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Vercel packs — Build Performance and Custom Domains Tls and the rest — against Vercel or your own agent with your own data.