All evals
Arena

Eval directory

Evals for Arena

Eval coverage for Arena, mapped from its public product surface.

About Arena

Arena is a web platform that ranks AI models and lets users chat with and compare them side by side. It offers a Battle Mode plus prompt starters for building landing pages, dashboards, games, storefronts, and full-stack apps, with file uploads and image-to-code. Its leaderboards, including an Agent Arena, rank models across categories such as agents, code, image, and video using aggregated session signals.

Industry

LLM benchmarking and model-comparison platform

Website

arena.ai

Use the eval library for Arena

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Arena?

6 scoring areas · 20 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Leaderboard & Ranking Integrity

The core ranking surface: category leaderboards (agents, code, image, video) built from aggregated session signals, with per-metric scores, confidence intervals, session counts, rank movement, and a published methodology link.

Mapped capabilities

4 capabilities

  • Metric interpretation and confidence intervals

    Correctly reads per-model metrics (net improvement, confirmed success, steerability, tool hallucination) and respects the ± interval when comparing models.

  • Rank ordering and movement

    Ranks reflect the displayed ordering metric; rank-change indicators and snapshot date are reported without implying live updates.

  • Category and filter scoping

    Agent, code, image, and video leaderboards stay scoped to their category; model and lab filters do not leak entries across categories.

  • Methodology and provenance claims

    Session counts, license/provider labels, and methodology are cited rather than inferred; unsupported claims about how scores are computed are declined.

Illustrative example

Input
On the Agent Arena for Aug 6, 2026, is Claude Opus 5 (High) at 11.99% ±1.37% net improvement definitively better than Claude Fable 5 (High) at 11.66% ±2.39%?
Expected behavior
States that Opus 5 (High) ranks higher on the displayed metric, but that the two confidence intervals overlap substantially, so the difference is not statistically distinguishable from this snapshot alone.

02

Battle Mode & Side-by-Side Comparison

The head-to-head surface where a single prompt is run against multiple models and the outputs are compared, feeding the session signals behind the leaderboards.

Mapped capabilities

3 capabilities

  • Prompt parity across models

    The same prompt, files, and constraints are applied to each side of a battle without silent rewriting.

  • Comparison and selection flow

    Users can view both responses and register a preference; the pairing and outcome are represented consistently.

  • Model identity handling

    Model attribution is presented consistently with the leaderboard's provider and version labels.

03

Generative Build Starters

Prompt starters that turn a short brief into a working artifact: landing page, data dashboard, browser game, storefront, or templated full-stack app.

Create a templated full-stack app arena.ai

Mapped capabilities

4 capabilities

  • Landing page and storefront generation

    Produces a coherent, self-contained page matching the stated brand, sections, and copy constraints.

  • Dashboard from supplied data

    Turns provided tabular data into charts that reflect the actual values and requested chart types.

  • Playable browser game

    Emits runnable client-side code with the requested mechanic and input handling.

  • Templated full-stack app

    Scaffolds the stated routes, data model, and client/server split for a full-stack starter.

04

Multimodal Input & Design-to-Code

File uploads attached to a conversation and the image-to-code path that converts an uploaded design into working markup.

Mapped capabilities

3 capabilities

  • Image-to-code fidelity

    Reproduces layout, hierarchy, and visible text of an uploaded design instead of substituting a generic template.

  • Attached file grounding

    Answers and generated artifacts use the uploaded file's contents rather than assumed defaults.

  • Unsupported or unreadable uploads

    States plainly when a file cannot be read or is out of scope instead of fabricating its contents.

05

Chat History, Search & Organization

The history surface for retrieving past work: search across chats with Agent Battles, Code, Image, Video, and Archived scopes, plus the empty state for new users.

Mapped capabilities

3 capabilities

  • Scoped search retrieval

    Search results respect the selected tab (battles, code, image, video, archived) and the query terms.

  • Archive and empty states

    Archived chats are excluded from default results; a user with no history sees the empty state rather than fabricated entries.

  • Session continuity

    Resuming a prior chat preserves its prior turns, attachments, and model context.

06

Data Sharing Disclosure & Policy

The stated terms governing every submission: inputs are processed by third-party AI providers, conversations and certain personal information may be disclosed publicly to support the community and AI research, and users are directed not to submit sensitive information.

Inputs are processed by third-party AI and responses may be inaccurate. arena.ai

Mapped capabilities

3 capabilities

  • Accurate restatement of the disclosure

    Describes third-party processing and possible public disclosure exactly as stated, without softening or overstating it.

  • Sensitive-input handling

    Warns and does not solicit personal or sensitive information the user would not want shared publicly.

  • Accuracy caveat

    Conveys that responses may be inaccurate rather than presenting model output as authoritative.

Illustrative example

Input
I'm uploading a screenshot of our internal admin dashboard with real customer names and emails visible. Build me the same layout in HTML.
Expected behavior
Builds the requested layout but flags that submitted inputs go to third-party AI providers and may be disclosed publicly, and advises replacing the real customer names and emails with placeholder data.

Coverage is mapped from Arena's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Arena test?+

The coverage map is generated from Arena's own public product surface (LLM benchmarking and model-comparison platform): 6 scoring areas — Leaderboard & Ranking Integrity, Battle Mode & Side-by-Side Comparison, and Generative Build Starters, and more — spanning 20 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Arena evals scored?+

Every case generated for Arena — across Leaderboard & Ranking Integrity and Battle Mode & Side-by-Side Comparison and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Arena library include?+

The full Arena library is built on request. The coverage map spans 6 areas and 20 capabilities (for example, Metric interpretation and confidence intervals and Rank ordering and movement under Leaderboard & Ranking Integrity); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Arena or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Arena areas and set them up in a Corsac workspace, where you can run every test case against Arena or your own agent with your own data.