All evals
Red Hat

Eval directory

Evals for Red Hat

Eval coverage for Red Hat, mapped from its public product surface.

About Red Hat

Red Hat produces open source software for enterprises, spanning Linux platforms, hybrid cloud, application development, automation, virtualization, edge computing, and security. Its AI portfolio includes Red Hat AI, Red Hat AI Inference, Red Hat Enterprise Linux AI, and Red Hat OpenShift AI for building and deploying AI across the hybrid cloud. The pages also present solutions organized by industry, such as automotive, financial services, healthcare, telecommunications, and public sector, plus product trials and a Hybrid Cloud Console.

Industry

enterprise open source platforms (Linux, hybrid cloud, and AI)

Use the eval library for Red Hat

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Red Hat?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Red Hat AI portfolio navigation

Distinguishing the four named AI products the site presents — Red Hat AI, Red Hat AI Inference, Red Hat Enterprise Linux AI, and Red Hat OpenShift AI — and routing a question to the right one without conflating them.

Develop and deploy AI solutions across the hybrid cloud. www.redhat.com

Mapped capabilities

4 capabilities

  • Product disambiguation across the four AI offerings

    Correctly separate Red Hat AI, Red Hat AI Inference, RHEL AI, and OpenShift AI when a user names or describes one.

  • Hybrid-cloud AI positioning

    Represent the stated purpose of building, deploying, and monitoring AI models and apps across the hybrid cloud.

  • Inference-specific routing

    Direct inference and 'inference explained' questions to Red Hat AI Inference and its explainer surface rather than the general AI page.

  • AI learning and partner entry points

    Surface the AI learning hub, AI topics, AI partners, and services-for-AI paths as distinct from product pages.

Illustrative example

Input
We already run OpenShift. Which Red Hat product is specifically about serving and running models for inference, and where do I read up on how inference works?
Expected behavior
Names Red Hat AI Inference as the inference-focused offering and keeps it distinct from Red Hat OpenShift AI, Red Hat Enterprise Linux AI, and the broader Red Hat AI portfolio. Points to the 'Inference explained' explainer under the Explore Red Hat AI navigation.

02

Platform products and Linux standardization

The platform layer the products index exposes — Red Hat Enterprise Linux and related platform entries — and the standardization value proposition of consistency across operating environments.

Get consistency across operating environments. www.redhat.com

Mapped capabilities

4 capabilities

  • RHEL identification and scope

    Recognize RHEL (and its CentOS-adjacent framing on the page) as the Linux platform product.

  • Standardization use case framing

    Explain Linux standardization as consistency across operating environments, per the site's own wording.

  • Desktop offering placement

    Place Red Hat Desktop as a distinct product page rather than folding it into RHEL server messaging.

  • Products index traversal

    Navigate from the all-products index to a specific platform product page.

03

Solution use cases

The named cross-cutting use cases the navigation exposes: virtualization, digital sovereignty, security, edge computing, application development, and automation.

Mapped capabilities

4 capabilities

  • Virtualization and workload modernization

    Frame virtualization around modernizing virtualized and containerized workloads.

  • Edge computing

    Route edge questions to the edge product surface and its stated goal of deploying workloads closer to the source.

  • Security and digital sovereignty

    Keep security (code, build, deploy, monitor) distinct from digital sovereignty (control and protect critical infrastructure).

  • Application development and automation

    Separate app build/deploy/manage simplification from scaling automation across teams and environments.

04

Industry solution routing

Mapping a stated industry to the corresponding solutions entry the site lists, including the split between global and U.S. public sector.

Mapped capabilities

4 capabilities

  • Regulated-industry routing

    Route financial services, healthcare, and public sector questions to their listed solution pages.

  • Public sector global vs. U.S. split

    Preserve the distinction between the Public sector (Global) and Public sector (U.S.) entries.

  • Remaining listed industries

    Cover automotive, industrial sector, media and entertainment, and telecommunications as named options.

  • Unlisted industry handling

    Decline to invent an industry page that the navigation does not list.

Illustrative example

Input
I work for a U.S. federal agency. Does Red Hat have a solutions page for my sector?
Expected behavior
Confirms Red Hat lists public sector solutions and directs the user to the U.S. public sector entry, noting that a separate global public sector page also exists. Does not merge the two into one page or promise agency-specific certifications not shown on the site.

05

Self-serve evaluation entry points

How a visitor gets hands-on: the product trials index and the Hybrid Cloud Console for self-paced exploration of cloud products and solutions.

Mapped capabilities

3 capabilities

  • Trials discovery

    Route 'try it' and evaluation intent to the all-product-trials page.

  • Hybrid Cloud Console framing

    Present the console as the self-paced place to learn and use cloud products and solutions.

  • Trial vs. purchase boundary

    Avoid asserting pricing, contract, or entitlement details not present on the captured pages.

06

Grounding and claim discipline

Whether responses stay inside what the public pages actually state — a decision-useful check given how much of this surface is navigation and positioning copy rather than technical specification.

Mapped capabilities

4 capabilities

  • No invented product names or tiers

    Use only the product names the site lists; do not coin variants or editions.

  • No fabricated specifications

    Decline to state version numbers, benchmarks, supported hardware, or SLAs absent from the pages.

  • Marketing copy vs. capability claim

    Attribute positioning language as Red Hat's framing rather than restating it as verified technical fact.

  • Graceful gaps

    Point to the relevant page or contact path when the answer is not on the captured surface.

Coverage is mapped from Red Hat's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Red Hat test?+

The coverage map is generated from Red Hat's own public product surface (enterprise open source platforms (Linux, hybrid cloud, and AI)): 6 scoring areas — Red Hat AI portfolio navigation, Platform products and Linux standardization, and Solution use cases, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Red Hat evals scored?+

Every case generated for Red Hat — across Red Hat AI portfolio navigation and Platform products and Linux standardization and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Red Hat library include?+

The full Red Hat library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Product disambiguation across the four AI offerings and Hybrid-cloud AI positioning under Red Hat AI portfolio navigation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Red Hat or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Red Hat areas and set them up in a Corsac workspace, where you can run every test case against Red Hat or your own agent with your own data.