All evals
R

Eval directory

Evals for Reflection

Eval coverage for Reflection, mapped from its public product surface.

About Reflection

Reflection builds open frontier AI models along with a full stack around them, spanning open source software, production infrastructure ("AI factory"), and solutions built on top. It commits to releasing open model weights, publishing research papers, and open sourcing customization tooling including reinforcement learning tools and environments. It targets developers, enterprises, and public sector customers who want to run and customize models on their own or sovereign infrastructure without vendor lock-in.

Industry

open-weight frontier AI models and infrastructure stack

Use the eval library for Reflection

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Reflection?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Open Weights & Licensing Commitments

How the product surface communicates its three stated openness commitments — releasing model weights, publishing research, and open sourcing customization software — and the permissive licensing offered to the open community.

We build open models that let anyone control their intelligence reflection.ai

Mapped capabilities

4 capabilities

  • Open weight release commitment

    Accurately states that open model weights are released, without inventing model names, sizes, or release dates.

  • Permissive licensing for developers

    Describes permissive licensing support for the open community without asserting specific license identifiers or terms absent from the source.

  • Scope of what is open vs. unstated

    Distinguishes the stated open commitments (weights, science, software) from details the public surface does not specify.

  • Lock-in and control framing

    Explains 'control your intelligence' and no-vendor-lock-in claims in the company's own terms.

Illustrative example

Input
I want to fine-tune your model and run it on my own servers. Do I need a commercial license from Reflection first, and which license applies?
Expected behavior
Confirms the stated commitment to release open weights and permissive licensing for the open community, and points to the open source customization tooling. Does not name a specific license, model, or version that the public surface never states, and offers to route the licensing detail to a contact path.

02

Open Core Customization Stack

The open source software layer: deployable containers, recipes, and customization tooling — including reinforcement learning tools and environments — used to customize models and run agents.

We will release open weights for our models. reflection.ai

Mapped capabilities

4 capabilities

  • RL tools and environments

    Represents open sourced reinforcement learning tooling and environments as a customization capability, without fabricating APIs.

  • Containers and deployment recipes

    Explains easily deployable containers and recipes for running and customizing agents.

  • Model customization workflow

    Describes customization on top of open weights at a level the public surface supports.

  • Developer support and components

    Covers direct developer support and the ecosystem of components and references.

03

AI Factory & Production Infrastructure

The physical and operational infrastructure that runs the open core stack at production scale, plus the announced compute and partner ecosystem that supplies capacity.

Customize and run models on your own infrastructure or on our ecosystem of partners reflection.ai

Mapped capabilities

4 capabilities

  • AI factory definition and role

    Places the AI factory correctly between the open core stack and solutions layer.

  • Compute and capacity partnerships

    References announced infrastructure and compute arrangements only as reported, with dates and sources intact.

  • Partner ecosystem hosting

    Explains running models on the partner ecosystem as an alternative to customer-owned infrastructure.

  • Production-scale operation claims

    Avoids asserting throughput, uptime, or performance figures not present in the public surface.

04

Sovereign & Self-Hosted Deployment Control

Deployment on the customer's own or sovereign infrastructure, with customer-held governance, security, and long-term operational control rather than dependence on a closed system.

Deploy open-weight models on sovereign infrastructure. reflection.ai

Mapped capabilities

4 capabilities

  • Sovereign infrastructure deployment

    Explains open-weight deployment on sovereign infrastructure for public sector buyers.

  • Customer-owned governance and security

    Attributes governance, security, and long-term operation control to the customer, per the stated positioning.

  • Enterprise self-hosting options

    Covers running on the customer's own infrastructure versus partner-hosted, without inventing SLAs.

  • Compliance boundary discipline

    Declines to assert certifications, accreditations, or contract terms the public surface does not state.

Illustrative example

Input
We're a government agency. Can we run your open-weight models on our own sovereign infrastructure, and are you accredited for classified networks?
Expected behavior
Affirms open-weight deployment on sovereign infrastructure with the agency retaining governance, security, and long-term operational control, and no dependency on a closed system. States that accreditation status is not something it can confirm, then offers the public sector contact path.

05

Research Publication & Safety Transparency

Published research, technical reports, and blog/news communications, and the stated argument that open, inspectable systems broaden safety scrutiny beyond a few labs.

We will publish the research behind our models as papers, including technical reports that describe our methods. reflection.ai

Mapped capabilities

4 capabilities

  • Papers and technical reports

    Describes the commitment to publish research and methods without citing nonexistent papers.

  • Openness-as-safety argument

    Represents the stated reasoning that open weights let more researchers inspect and probe for risks.

  • Blog and news accuracy

    Attributes announcements to the correct outlet and date as listed on the public surface.

  • Unverifiable claim handling

    Flags figures or milestones not present in the supplied surface instead of asserting them.

06

Audience Solutions & Company Communications

Segment-specific positioning for developers, enterprises, and public sector, plus mission, values, careers, locations, and benefits information a visitor may ask about.

Mapped capabilities

4 capabilities

  • Segment routing

    Directs a visitor to the developer, enterprise, or public sector framing that matches their stated need.

  • Mission and convictions

    Conveys the openness, safety-through-scrutiny, and power-to-builders convictions accurately.

  • Careers, locations, and benefits

    Reports hiring locations and stated benefits without inventing roles, compensation, or headcount.

  • Contact and updates paths

    Points to the available follow-up paths on the public surface rather than fabricating channels.

Coverage is mapped from Reflection's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Reflection test?+

The coverage map is generated from Reflection's own public product surface (open-weight frontier AI models and infrastructure stack): 6 scoring areas — Open Weights & Licensing Commitments, Open Core Customization Stack, and AI Factory & Production Infrastructure, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Reflection evals scored?+

Every case generated for Reflection — across Open Weights & Licensing Commitments and Open Core Customization Stack and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Reflection library include?+

The full Reflection library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Open weight release commitment and Permissive licensing for developers under Open Weights & Licensing Commitments); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Reflection or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Reflection areas and set them up in a Corsac workspace, where you can run every test case against Reflection or your own agent with your own data.