All evals
RadixArk

Eval directory

Evals for RadixArk

Eval coverage for RadixArk, mapped from its public product surface.

About RadixArk

RadixArk is an infrastructure-first deep-tech company building large-scale inference and training systems for the AI community. It builds on SGLang for inference serving and Miles, its open-source framework for large-scale RL post-training, and layers managed infrastructure and tooling on top for developers, startups, enterprises, and research labs. A hosted platform at platform.radixark.com lets registered users run open-source models and receive outputs from their prompts and uploaded content.

Industry

AI inference and training infrastructure (open-source serving + RL post-training platform)

Use the eval library for RadixArk

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for RadixArk?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Hosted Model Platform Workflow

The registered-user experience at platform.radixark.com: accessing open source models and receiving outputs from queries, prompts, text, images, and other uploaded or submitted content.

the fastest, most flexible open engine for serving modern models www.radixark.com

Mapped capabilities

4 capabilities

  • Open-source model access and selection

    Explaining that the platform exposes different open source models users can access and run, without asserting a specific unlisted model roster.

  • Prompt and query response flow

    Describing how a user submits queries, inputs, or prompts and receives model outputs through the Platform.

  • Uploaded content as model input

    Handling of text, documents, images, photos, videos, and audio files submitted through the Services as inputs.

  • Registration gate for platform use

    Stating that certain portions of the Platform require registering an account with name, email, and other requested information.

Illustrative example

Input
I don't have an account yet. Can I upload a PDF to your platform right now and get one of your models to summarize it?
Expected behavior
Explains that registration with name, email, and other requested information is required for certain portions of the Platform, and that the Platform runs open source models over submitted prompts and uploaded content. Does not claim anonymous or immediate access.

02

Inference Serving on SGLang

RadixArk's inference foundation: SGLang as the open engine for serving modern models, positioned as the performance and reliability base for production AI applications.

RadixArk provides a platform where users can access and use different open source models and receive outputs www.radixark.com

Mapped capabilities

4 capabilities

  • SGLang as the serving foundation

    Attributing inference serving to SGLang and explaining its role as the engine underneath RadixArk's inference offering.

  • Continued investment posture

    Communicating RadixArk's stated commitment to keep investing in SGLang rather than forking away from it.

  • Production reliability framing

    Explaining the claim that SGLang serves as the performance and reliability foundation for production AI applications, without inventing benchmarks.

  • Openness of the engine

    Distinguishing the open engine from the managed layer RadixArk ships on top of it.

03

RL Post-Training with Miles

Miles, RadixArk's open-source framework for large-scale reinforcement learning post-training, described as bringing serving-engine rigor to RL.

our open-source framework for large-scale post-training www.radixark.com

Mapped capabilities

4 capabilities

  • Miles scope and purpose

    Identifying Miles as RadixArk's open-source framework for large-scale RL post-training.

  • Training vs. inference separation

    Correctly routing a question to Miles (RL training) or SGLang (inference serving) rather than conflating the two.

  • Open-source availability

    Stating that Miles is open source and part of the cores RadixArk builds on and ships around.

  • Rigor-to-RL positioning

    Explaining the stated analogy that Miles brings to RL the rigor modern serving engines brought to inference.

04

Managed Infrastructure and Audience Fit

The managed infrastructure and tooling layered on the open cores, aimed at individual developers, startups, enterprises, and research labs.

make frontier-level AI infrastructure open and accessible to everyone www.radixark.com

Mapped capabilities

4 capabilities

  • Managed layer over open cores

    Explaining that RadixArk ships managed infrastructure and tooling on top of SGLang and Miles.

  • Audience routing

    Matching an inquiry to the four stated audiences — individual developers, startups, enterprises, research labs — without promising unlisted tiers.

  • Rebuild-avoidance value proposition

    Articulating the stated problem of every new AI lab rebuilding schedulers, compilers, serving engines, and training pipelines.

  • Engineering philosophy claims

    Representing the first-principles, elegance-and-throughput, frontier-scale-reliability stance as stated positioning.

05

Accounts, Terms, and Billing

Agreement scope and commercial mechanics from the Terms of Service: acceptance, entity binding, account registration, arbitration, and payment handling.

THEY CONTAIN AN AGREEMENT TO ARBITRATE AND OTHER IMPORTANT INFORMATION REGARDING YOUR LEGAL RIGHTS www.radixark.com

Mapped capabilities

4 capabilities

  • Agreement scope and acceptance

    That the Agreement covers radixark.com, platform.radixark.com, and linked products, and that access or use signifies agreement.

  • Binding an organization

    The representation required when accepting on behalf of a company, business, or other legal entity.

  • Arbitration and legal rights notice

    Surfacing that the Terms contain an agreement to arbitrate and other information about legal rights, remedies, and obligations.

  • Payment and billing information

    That payment method, billing address, bank account, and transaction history are collected and processed by a third-party payment processor.

06

Privacy and Data-Handling Scope

What the Privacy Policy covers across websites, the AI infrastructure platform, marketing and sales, and job applications — and the categories it explicitly excludes.

does not apply to our handling of personal information that we process on behalf of our business customers www.radixark.com

Mapped capabilities

4 capabilities

  • Business-customer carve-out

    That the Privacy Policy does not apply to personal information processed on behalf of business customers, which is governed by customer agreements.

  • Categories of information collected

    Account and contact information, user content, company and workspace information, payment information, feedback, and support communications.

  • Covered surfaces

    That the Services span both websites, the AI infrastructure platform, marketing and sales initiatives, and job applications.

  • Workspace and team context

    Handling of organization name, workspace identifiers, and team member roles collected in connection with use of the Services.

Illustrative example

Input
We're onboarding as a business customer. Does your Privacy Policy govern the personal data our end users push through your infrastructure platform?
Expected behavior
States that the Privacy Policy does not apply to personal information RadixArk processes on behalf of business customers, and that such processing is governed instead by the agreements RadixArk has with those business customers.

Coverage is mapped from RadixArk's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for RadixArk test?+

The coverage map is generated from RadixArk's own public product surface (AI inference and training infrastructure (open-source serving + RL post-training platform)): 6 scoring areas — Hosted Model Platform Workflow, Inference Serving on SGLang, and RL Post-Training with Miles, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the RadixArk evals scored?+

Every case generated for RadixArk — across Hosted Model Platform Workflow and Inference Serving on SGLang and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the RadixArk library include?+

The full RadixArk library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Open-source model access and selection and Prompt and query response flow under Hosted Model Platform Workflow); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against RadixArk or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped RadixArk areas and set them up in a Corsac workspace, where you can run every test case against RadixArk or your own agent with your own data.