All evals
Runway

Eval directory

Evals for Runway

Eval coverage for Runway, mapped from its public product surface.

About Runway

Runway builds foundational "Real-World Intelligence" general world models and sells three product surfaces on top of them: Runway Creative (a cloud creative suite for generating and editing video, images and audio), Runway Dev (an API/media platform for developers to combine models into pipelines and workflows), and Runway Robotics (policy inference, offline policy evaluation and synthetic data generation powered by GWM-1). Consumer plans range from a free tier to paid credit-based subscriptions, with tools including Gen-4.5, Aleph 2.0 in-context video editing, Runway Agent for marketing campaigns, and real-time conversational Characters. A Builders Program offers Seed–Series C startups free API credits and early access to the Characters API.

Industry

generative AI video, image and world-model platform

Website

runway.com

Use the eval library for Runway

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Runway?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Creative Generation & Model Selection

The all-in-one cloud creative suite: generating video, image and audio across a large catalog of first- and third-party models, and routing a request to the right model and tool.

Mapped capabilities

4 capabilities

  • Text/image/video/audio-to-video generation

    Gen-4.5, Gen-4 Turbo, Seedance and Kling-class generation from mixed input modalities

  • Model catalog routing and availability

    Choosing among Gen-4.5, Nano Banana Pro, Seedream, FLUX, Eleven and others; handling models marked coming soon

  • Multi-shot and structured outputs

    Multi-Shot Video App: one prompt to a coherent multi-shot sequence, storyboards and animatics

  • Image generation and upscaling

    1080p/2K image generation, 4K upscaling and Topaz AI upscale as plan-gated steps

02

In-Context Video Editing (Aleph 2.0 & Edit Studio)

Editing an existing asset rather than generating a new one: change only what was asked, propagate the change across the clip, and preview before committing credits.

Mapped capabilities

4 capabilities

  • Frame-edit propagation

    Edit one frame and carry the look through the rest of the video

  • Scoped element change with preservation

    Change a product color, hairstyle or garment while background, lighting and other details stay fixed

  • Multi-shot and clip-length limits

    Applying edits across cuts; 30s at 1080p ceiling

  • Edit Studio pre-generation preview

    Image preview of the intended edit to reduce iteration count before generating

Illustrative example

Input
Use Aleph 2.0 to recolor only the jacket in my 45-second 4K interview clip, leaving the background and lighting untouched.
Expected behavior
Confirms Aleph 2.0 changes a scoped element while preserving background and lighting, but flags that it supports clips up to 30 seconds at 1080p, so the 45-second 4K source must be trimmed or split first.

03

Runway Agent for Marketing Workflows

Conversational, end-to-end marketing work: concepting a brief, diagnosing underperforming creative, generating test variations, and scaling a winner across formats, channels and markets.

Mapped capabilities

4 capabilities

  • Brief-to-campaign asset production

    Positioning plus deliverables for a stated product, audience and launch

  • Campaign data analysis

    Reading shared campaign performance data to identify what is driving results

  • Variant and hook experimentation

    Fresh hooks and A/B variations targeted at Meta, TikTok and YouTube

  • Format resizing and localization

    9:16, 16:9 and 1:1 sizing; translating a top performer for each market

04

Characters: Real-Time Conversational Agents

Persistent AI personas powered by GWM-1 with a defined appearance, voice and personality, created once and reused across sessions and deployment surfaces.

Mapped capabilities

4 capabilities

  • Persona definition and persistence

    Stable appearance, voice and personality reused across every session

  • Real-time conversational behavior

    Natural back-and-forth video conversation as an interactive agent

  • Characters API access and deployment

    Create once, deploy anywhere via the Runway API; access gated to Builders and Enterprise

  • Custom voice and lip sync

    Custom voices for lip sync and text to speech as a Pro-tier capability

05

Runway Dev: API, Pipelines & Integrations

The developer media platform for combining multiple models, modalities and tasks into pipelines, and triggering them from custom endpoints or external agents.

Unlock Tier 5 access — Runway's highest API usage tier — for maximum throughput runway.com

Mapped capabilities

4 capabilities

  • Multi-model pipeline composition

    Chaining models, modalities and tasks into one workflow

  • Custom Workflow endpoints

    Triggering a saved Workflow via its own endpoint

  • MCP and external agent integration

    Connecting Runway to Claude, ChatGPT and similar agents through Runway MCP

  • Usage tiers and rate limits

    API key issuance, tier-based throughput, Tier 5 as the highest usage tier

06

Runway Robotics & GWM-1 Toolkit

Bringing the General World Model into a robot learning pipeline across four distinct surfaces, fitted to the customer's hardware without specialized compute to start.

Runway is building foundational Real-World Intelligence that can understand, simulate and act in the world. runway.com

Mapped capabilities

4 capabilities

  • Policy inference from observations

    GWM-1 predicting actions from live camera observations, fine-tuned per hardware, environment and task

  • Offline policy evaluation

    Simulating rollouts from submitted action sequences and observations to surface failure modes pre-deployment

  • Synthetic data augmentation

    Transforming existing trajectories into new environments, lighting and object configurations

  • Video model licensing

    GWM-1 as a diffusion backbone trained and deployed on customer infrastructure

Coverage is mapped from Runway's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Runway test?+

The coverage map is generated from Runway's own public product surface (generative AI video, image and world-model platform): 6 scoring areas — Creative Generation & Model Selection, In-Context Video Editing (Aleph 2.0 & Edit Studio), and Runway Agent for Marketing Workflows, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Runway evals scored?+

Every case generated for Runway — across Creative Generation & Model Selection and In-Context Video Editing (Aleph 2.0 & Edit Studio) and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Runway library include?+

The full Runway library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Text/image/video/audio-to-video generation and Model catalog routing and availability under Creative Generation & Model Selection); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Runway or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Runway areas and set them up in a Corsac workspace, where you can run every test case against Runway or your own agent with your own data.