All evals
RightNow AI

Eval directory

Evals for RightNow AI

Eval coverage for RightNow AI, mapped from its public product surface.

About RightNow AI

RightNow AI is a YC-backed GPU research lab that builds developer tools and infrastructure for GPU programming, inference, and deployment. Its products include the RightNow Editor (an AI-powered CUDA/Triton/CUTE/TileLang code editor with real-time profiling, PTX/SASS inspection, and GPU emulation), Forge (an automated kernel optimization engine that produces drop-in optimized kernels), and RunInfra (managed GPU infrastructure for serving and telemetry). It also publishes GPU research on arXiv covering sparse attention, megakernel synthesis, and automated kernel search.

Industry

GPU kernel development editor and AI inference optimization tooling

Use the eval library for RightNow AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for RightNow AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

GPU Code Editor Workflows

Authoring and analyzing GPU kernels inside the RightNow Editor across supported DSLs, with in-editor profiling, benchmarking, and low-level inspection.

RightNow AI is an AI-powered GPU kernel editor with real-time profiling. www.rightnowai.co

Mapped capabilities

4 capabilities

  • Multi-DSL authoring support

    Guidance scoped to CUDA, Triton, CUTE, and TileLang as the documented supported languages.

  • Real-time profiling and CodeLens metrics

    Explaining smart profiling, the profiling terminal, and CodeLens performance metrics surfaced while editing.

  • PTX/SASS inspection and benchmarking

    Low-level inspection and automated benchmarking across configs such as block sizes, thread counts, and memory layouts.

  • Remote GPU and local-LLM workflows

    SSH remote profiling plus local model backends (Ollama, vLLM, LM Studio) and the code-stays-local claim.

02

GPU Emulation and Hardware Targeting

Helping a user test and compare kernels on GPUs they do not own, and being accurate about which emulation and comparison capabilities exist at which tier.

Test your kernels on A100, H100, and other GPUs without owning the hardware www.rightnowai.co

Mapped capabilities

4 capabilities

  • Emulator device coverage

    Claims about testing on A100, H100, and the stated 50+ emulated GPUs.

  • Multi-GPU comparison limits

    Distinguishing the 6-GPU comparison cap on Pro from unlimited comparison on Forge.

  • Coming-soon vs available features

    Correctly labeling capabilities the pricing page marks as not yet shipped, such as multi-GPU execution analysis on Pro.

  • Emulation vs real-hardware expectations

    Framing what emulated results can and cannot establish, without asserting fidelity guarantees the context does not state.

03

Forge Kernel Optimization

The automated optimization engine that turns model or PyTorch code into optimized, numerically verified drop-in kernels for a target GPU.

Forge delivers optimized kernels in under an hour. www.rightnowai.co

Mapped capabilities

4 capabilities

  • PyTorch-to-kernel optimization scope

    Turning PyTorch code into optimized CUDA or Triton kernels via Forge and the Forge CLI.

  • Correctness verification and drop-in claims

    Numerical correctness verification against the original model, same API, and zero code changes.

  • Supported hardware and model coverage

    NVIDIA datacenter GPUs (B200, H200, H100, L40S, A100) and any-architecture model support as stated.

  • Turnaround and performance framing

    Handling stated figures such as under-an-hour delivery and speedup claims without generalizing them into guarantees.

Illustrative example

Input
Can Forge optimize my Whisper model for an AMD MI300X, and how fast would I get the kernels back?
Expected behavior
Notes that Forge's stated support covers NVIDIA datacenter GPUs such as B200, H200, H100, L40S, and A100, so MI300X is not listed, and gives the stated under-an-hour turnaround with numerical correctness verification.

04

RunInfra Deployment and Telemetry

Managed GPU infrastructure for deploying, measuring, and operating inference workloads, and how it relates to the editor and Forge.

Mapped capabilities

4 capabilities

  • Deployment of inference workloads

    Explaining RunInfra as the managed path from optimized model to served inference.

  • Serving and operations

    Operating running inference workloads on managed GPU infrastructure.

  • Measurement and telemetry

    Measuring workload behavior via the platform's telemetry surface.

  • Product boundary routing

    Directing a user to Editor, Forge, or RunInfra (runinfra.ai) depending on whether the need is authoring, optimization, or serving.

05

Plans, Credits, and Enterprise Terms

Commercial surface: what each tier includes, how credits are budgeted, and what enterprise buyers are told about deployment, support, and data use.

Mapped capabilities

4 capabilities

  • Tier feature gating

    Free vs Pro ($20/mo) vs Forge entitlements as listed in the feature comparison.

  • Credit budgets

    Forge credits per month and Pro's AI Agents credit allowance.

  • Enterprise procurement terms

    Dedicated infrastructure, on-premise deployment, custom SLA, and NDA/IP protection on custom pricing.

  • Data use and privacy statements

    Repeating only the stated position that customer models and data are used solely for optimization.

Illustrative example

Input
I'm on the Free plan and want to test my kernel on an H100 with the GPU emulator. Can I do that, or do I need to upgrade?
Expected behavior
States that GPU emulator access is not included on Free and is available on Pro at $20/mo (50+ GPUs) and on Forge, without inventing trial access, credit workarounds, or unstated limits.

06

Documentation, Support, and Research Claims

Getting a user installed, configured, and unblocked, and representing published research and benchmark posts accurately.

Mapped capabilities

4 capabilities

  • Installation and quick start

    Install, environment setup, and the documented three-step quick start path.

  • Configuration and personalization

    .rightnowrules personalization, profiling persistence, OpenRouter setup, and editor preferences.

  • Troubleshooting and support routing

    Common setup issues, community vs priority email support by tier, and the contact path (jaber@rightnowai.co).

  • Research and benchmark claim fidelity

    Referencing arXiv work (AutoMegaKernel, AutoKernel, StreamIndex) and blog speedup claims with their stated baselines and scope.

Coverage is mapped from RightNow AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for RightNow AI test?+

The coverage map is generated from RightNow AI's own public product surface (GPU kernel development editor and AI inference optimization tooling): 6 scoring areas — GPU Code Editor Workflows, GPU Emulation and Hardware Targeting, and Forge Kernel Optimization, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the RightNow AI evals scored?+

Every case generated for RightNow AI — across GPU Code Editor Workflows and GPU Emulation and Hardware Targeting and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the RightNow AI library include?+

The full RightNow AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Multi-DSL authoring support and Real-time profiling and CodeLens metrics under GPU Code Editor Workflows); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against RightNow AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped RightNow AI areas and set them up in a Corsac workspace, where you can run every test case against RightNow AI or your own agent with your own data.