All evals
R

Eval directory · AI Platform

Evals for Reka

Eval coverage for Reka, mapped from its public product surface.

About Reka

Reka builds multimodal AI for the physical world, spanning research models, an inference engine and API, video processing infrastructure, and training-data infrastructure. Its flagship product surface, Reka Vision, lets organizations search, reason over, monitor, and clip video and image content at scale. Related offerings include Reka Edge for edge intelligence and Claru for frontier AI training data.

Industry

multimodal video intelligence and inference infrastructure

Headquarters

Sunnyvale, California

Use the eval library for Reka

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Reka?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Video Understanding & Indexing

Turning raw video and images into structured, searchable representations — the substrate everything else in Reka Vision depends on.

Reka Vision is a next-generation multimodal AI system built to interpret, search, and reason across video and image content at scale. reka.ai

Mapped capabilities

4 capabilities

  • Object, action, scene, and event recognition

    Correct identification and labeling of what appears and what happens in a frame or shot.

  • Temporal segmentation across long videos

    Splitting multi-hour content into coherent segments with defensible boundaries.

  • Multimodal embedding and index construction

    Producing representations that support downstream retrieval across frames and modalities.

  • Structured output fidelity

    Tags and metadata that are internally consistent and tied to specific timestamps.

02

Semantic Video Search & Retrieval

Natural-language search over video and image libraries using meaning rather than filenames or manual metadata.

Scalable infrastructure to tag, reason over, search, and clip large volumes of video data. reka.ai

Mapped capabilities

4 capabilities

  • Natural-language query interpretation

    Mapping conversational queries to the intended visual content, including negations and constraints.

  • Cross-frame and cross-scene retrieval

    Finding matches that span shots rather than single-frame keyword hits.

  • Long-horizon event discovery

    Surfacing events that unfold over extended durations in large archives.

  • No-match and low-confidence handling

    Saying nothing matched instead of returning loosely related footage.

Illustrative example

Input
Search this 3-hour loading-dock recording for moments where someone drops a box, and return the matching moments.
Expected behavior
Returns matching moments as timestamped segments within the recording, ranked by relevance. If no moment matches, it states that no match was found rather than returning unrelated footage.

03

Highlight, Clip & Caption Generation

Automatically producing highlights, summaries, clips, and captions from long-form visual content.

Automatically generates highlights, summaries, and clips from visual content. reka.ai

Mapped capabilities

4 capabilities

  • Long-to-short clip extraction

    Selecting clip boundaries that are self-contained and within source bounds.

  • Intent-conditioned highlight detection

    Honoring a stated user intent or topic when choosing what to highlight.

  • ASR-based captioning

    Speech-derived captions aligned to the clip they accompany.

  • Summarization of visual content

    Summaries that reflect what is actually shown rather than inferred narrative.

Illustrative example

Input
From this 45-minute interview, generate three 30-second highlight clips about the guest's product launch, with captions.
Expected behavior
Produces exactly three clips of about thirty seconds each, drawn from within the source and tied to the product-launch discussion, each accompanied by speech-derived captions covering its audio.

04

Visual Reasoning, Q&A & Monitoring

Answering questions over video and images and monitoring content at scale, beyond recognition-level labeling.

Mapped capabilities

4 capabilities

  • Grounded question answering over video

    Answers tied to identifiable moments in the source rather than unsupported assertions.

  • Reasoning about physical-world events

    Inferences about actions, causes, and sequence consistent with the observed footage.

  • Monitoring and alerting behavior

    Flagging conditions of interest across incoming content at scale.

  • Abstention on unobservable detail

    Declining to answer when the requested detail is not visible in the material.

05

Inference, Integration & Edge Deployment

The inference engine and API plus the documented integration and deployment surfaces: API, MCP, the n8n community node, OpenRouter, and Reka Edge.

Inference engine and API for multimodal AI. Built for speed, scale, and enterprise reliability. reka.ai

Mapped capabilities

4 capabilities

  • API and MCP access paths

    Consistent behavior when the same capability is reached via API, MCP, or the app.

  • Model switching without code changes

    Swapping models (e.g. Reka Edge via OpenRouter) while callers stay unchanged.

  • Edge intelligence deployment

    Capability and constraint differences when running at the edge versus hosted.

  • Workflow-tool integration

    Video and image analysis invoked from an external automation workflow.

06

Training Data & World-Model Evaluation

Claru's training-data infrastructure and Reka Labs' published datasets and benchmarks for world models and physical realism.

Switch Models, Zero Code Changes: Reka Edge Now Available on OpenRouter reka.ai

Mapped capabilities

4 capabilities

  • Egocentric and manipulation data pipelines

    Collection and description quality for household, gameplay, and robotics footage.

  • World-model data preparation

    Preparing trajectories and footage for vision-language-action and omni world models.

  • Physical realism evaluation

    Detecting and attributing physics violations in synthetic or generated video.

  • Expert human judgment at scale

    Quality and consistency of human-provided labels and judgments.

Coverage is mapped from Reka's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Reka test?+

The coverage map is generated from Reka's own public product surface (multimodal video intelligence and inference infrastructure): 6 scoring areas — Video Understanding & Indexing, Semantic Video Search & Retrieval, and Highlight, Clip & Caption Generation, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Reka evals scored?+

Every case generated for Reka — across Video Understanding & Indexing and Semantic Video Search & Retrieval and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Reka library include?+

The full Reka library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Object, action, scene, and event recognition and Temporal segmentation across long videos under Video Understanding & Indexing); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Reka or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Reka areas and set them up in a Corsac workspace, where you can run every test case against Reka or your own agent with your own data.