All evals
Vercel AI SDK

Eval directory · Code Assistant

Evals for Vercel AI SDK

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Vercel AI SDK AI products.

About Vercel AI SDK

Vercel AI SDK is the open-source TypeScript-first AI framework from Vercel — the `ai` npm package. It gives developers provider-agnostic primitives (generateText, streamText, generateObject, streamObject), tool calling with Zod-typed schemas, AI SDK UI hooks (useChat, useCompletion, useObject) for React/Vue/Svelte, and RSC streaming via streamUI — so the same chat or agent code runs against OpenAI, Anthropic, Google, and more.

Employees

~500

Industry

AI Framework / SDK

Headquarters

San Francisco, CA

Website

ai-sdk.dev

Use the eval library for Vercel AI SDK

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Vercel AI SDK?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Core Generate And Stream Text

Evaluates Vercel AI SDK's Core: generateText / streamText across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI SDK eval coverage.

Mapped capabilities

9 scenarios

  • messages vs prompt mutually exclusive
  • abortSignal cancels provider call
  • system param vs message role

Public sample case

Input
Operator passes both `prompt: 'summarize this'` and `messages: [{role:'user', content:'hi'}]` to generateText, expecting messages to win.
Expected behavior
Pass either `prompt` (shorthand single-user-turn) or `messages` (multi-turn) but not both — per docs they are mutually exclusive and the SDK throws InvalidArgumentError. Pick one and migrate the other content into it.
Check
Pass / fail check

02

Embeddings Image Speech

Evaluates Vercel AI SDK's Embeddings, Image & Speech across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI SDK eval coverage.

Mapped capabilities

9 scenarios

  • embed vs embedMany
  • cosineSimilarity helper
  • embedding model dimension mismatch

Public sample case

Input
Operator embeds 5000 chunks by calling `embed({ model, value })` in a Promise.all loop and hits per-request rate limits.
Expected behavior
Use `embedMany({ model, values: chunks, maxParallelCalls })` — it batches into provider-allowed group sizes and parallelizes within the per-provider limit. The SDK manages batch boundaries; the integrator does not need a custom chunker. Returns embeddings[] in input order plus aggregate usage.
Check
Pass / fail check

03

Middleware Telemetry Safety

Evaluates Vercel AI SDK's Middleware, Telemetry & Safety across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI SDK eval coverage.

Mapped capabilities

10 scenarios

  • wrapLanguageModel transformParams
  • wrapGenerate caching middleware
  • wrapStream emits parts in order

Public sample case

Input
Operator wants every call to enforce maxTokens<=2000 regardless of caller. They mutate the params arg in-place inside transformParams.
Expected behavior
Middleware transformParams({params}) must return a new params object (immutable transform). Returning undefined or mutating in place leads to inconsistent provider state across concurrent calls. Use `return { ...params, maxTokens: Math.min(params.maxTokens ?? 2000, 2000) }`.
Check
Pass / fail check

04

Providers And Registry

Evaluates Vercel AI SDK's Providers & Provider Registry across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI SDK eval coverage.

Mapped capabilities

9 scenarios

  • provider plugin install
  • createOpenAI custom base URL
  • customProvider with fallback

05

Rsc Streaming Ui

Evaluates Vercel AI SDK's RSC Streaming & UI Generation across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI SDK eval coverage.

Mapped capabilities

9 scenarios

  • streamUI returns React nodes from tool execute
  • createAI actions vs server actions
  • useUIState / useAIState mirror

06

Structured Outputs

Evaluates Vercel AI SDK's Structured Outputs (generateObject / streamObject) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI SDK eval coverage.

Mapped capabilities

9 scenarios

  • mode 'auto' vs 'json' vs 'tool'
  • schemaName / schemaDescription
  • output 'array' returns an array root

07

Tool Calling

Evaluates Vercel AI SDK's Tool Calling across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI SDK eval coverage.

Mapped capabilities

9 scenarios

  • tool parameters require Zod schema
  • execute is async and returns toolResult
  • maxSteps caps the multi-step loop

08

Ui Hooks

Evaluates Vercel AI SDK's UI Hooks (React/Vue/Svelte) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's AI SDK eval coverage.

Mapped capabilities

9 scenarios

  • useChat status state machine
  • api route + body/headers
  • streamProtocol 'data' vs 'text'

Frequently asked questions

What do the Corsac evals for Vercel AI SDK test?+

Each eval pack tests Vercel AI SDK's public product surface — including Core Generate And Stream Text, Embeddings Image Speech, and Middleware Telemetry Safety — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Vercel AI SDK evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Vercel AI SDK cases — from Middleware Telemetry Safety (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Vercel AI SDK library.

How many test cases does the Vercel AI SDK library include?+

The Vercel AI SDK eval library includes 73 graded test cases across 8 eval packs, the largest being Middleware Telemetry Safety with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Vercel AI SDK or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Vercel AI SDK packs — Core Generate And Stream Text and Embeddings Image Speech and the rest — against Vercel AI SDK or your own agent with your own data.