All evals
Nous Research

Eval directory · Search & Knowledge

Evals for Nous Research

Eval coverage for Nous Research, mapped from its public product surface.

About Nous Research

Nous Research trains open-source language models and builds infrastructure for distributed model training. Its main lines are the Hermes series of steerable/reasoning models plus Hermes Agent, an autonomous server-resident agent, and Psyche, a decentralized training network built on DisTrO/DeMo and coordinated via Solana. It also publishes research, datasets, and RL tooling such as Atropos and NousCoder-14B.

Industry

open-source language models and distributed AI training infrastructure

Use the eval library for Nous Research

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Related in Search & Knowledge

All evals →

More Search & Knowledge eval libraries

Coverage map

What would you measure for Nous Research?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Hermes model line and steerability

Questions about the Hermes series as released: variants and parameter sizes, hybrid reasoning versus chat modes, extended context claims, and how steerability is described. Covers whether the surface answers model-selection questions accurately and separates released facts from speculation.

Hermes 4.3 was trained with an extended context length (up to 512K) nousresearch.com

Mapped capabilities

4 capabilities

  • Variant and size disambiguation

    Distinguishing Hermes 4 405B / 70B / 14B, Hermes-4.3-Seed-36B, DeepHermes-3-8B, and Hermes-3-Llama-3.2-3B by size and stated purpose.

  • Reasoning vs chat mode behavior

    Explaining hybrid-mode reasoning and deep-reasoning/chat mode toggles as documented, without inventing configuration flags.

  • Context length and local-inference fit

    Handling the 512K extended-context claim for Hermes 4.3 and which variants are positioned for local hardware.

  • Steerability framing

    Describing Hermes fine-tunes as highly steerable in the terms the company uses, without overclaiming safety or alignment properties.

Illustrative example

Input
I have one 24GB GPU. Which Hermes 4 model should I run, and roughly how large is it?
Expected behavior
Names Hermes-4-14B as the small, dense variant positioned for local inference and gives its listed size of 28 (GB per the releases table). Does not recommend the 70B or 405B variants for this hardware, and does not invent quantization or throughput numbers.

02

Hermes Agent

The autonomous agent released 02/25/26 that lives on a server, persists what it learns, and improves over long runs. Covers how the agent's autonomy, memory, and enterprise deployment story are represented to a technical buyer.

our open-source self-improving AI agent with a built-in learning loop nousresearch.com

Mapped capabilities

4 capabilities

  • Autonomy and residency model

    Explaining what 'lives on your server' implies for hosting, run duration, and operator control as described publicly.

  • Memory and self-improvement loop

    Representing the built-in learning loop and persistence of learned material without asserting unpublished mechanisms.

  • Enterprise deployment surface

    Hermes Agent Enterprise adaptation inside customer environments, per the forward-deployed engineering role description.

  • Research lineage

    Connecting the agent to the published 'Taming LLMs with Sequential Monte Carlo' research where the context supports it.

03

Psyche distributed training network

The decentralized training network: DisTrO/DeMo bandwidth reduction, Solana-based coordination, live training runs, and the two-stage roadmap. Covers accuracy about how participation and coordination actually work.

We train world-class open source language models and build infrastructure to coordinate distributed, unbiased training. nousresearch.com

Mapped capabilities

4 capabilities

  • DisTrO/DeMo bandwidth claims

    Stating the orders-of-magnitude data-transfer reduction and its lineage from DeMo without inflating benchmark numbers.

  • Solana coordination layer

    Explaining what the blockchain coordinates (fault tolerance, censorship resistance) versus what it does not do.

  • Contributing compute and live runs

    Directing an operator to live training runs and describing heterogeneous-hardware participation as documented.

  • Roadmap staging

    Distinguishing stage 1 cooperative training (permissioned testnet) from later accessible inference and full decentralization.

Illustrative example

Input
Does Psyche run the model training itself on Solana?
Expected behavior
Answers no: Solana is the coordination layer that makes the network fault-tolerant and censorship-resistant, while training runs on distributed, heterogeneous GPUs across the internet using DisTrO/DeMo to cut data transfer. Does not claim on-chain computation, token mechanics, or rewards not present in the source.

04

Open artifacts and release provenance

The releases index spanning models, datasets, papers, frameworks, and simulators with dates and sizes. Covers whether the surface can answer 'what was released, when, and under what artifact type' without conflation.

Mapped capabilities

4 capabilities

  • Release type and date accuracy

    Correctly typing an entry as MODEL, PAPER, DATASET, FRAMEWORK, AGENT, or SIMULATOR and citing its listed date.

  • Base-model attribution

    Attributing post-trained models to their bases (NousCoder-14B on Qwen3-14B, Hermes 4 on Llama 3.1) accurately.

  • Dataset and weight availability

    Pointing to Hugging Face weights, the Hermes 3 dataset, and open RL environments as the published distribution channels.

  • Partnership and provenance edge cases

    Handling entries with external partners or unusual provenance, such as Nomos 1 with Hillclimb and Psyche-trained checkpoints.

05

Research communication and RL tooling

Blog and framework surface: Atropos RL environments, tinker-atropos, contrastive neuron attribution, Lighthouse Attention, token superposition. Covers explaining technical results at the level the posts support, with author attribution.

Token Superposition Training (TST) delivers 2-3x wall-clock pretraining speedups at fixed FLOPs. nousresearch.com

Mapped capabilities

4 capabilities

  • Atropos and tinker-atropos scope

    Describing Atropos as an RL environments framework and tinker-atropos as an integration layer with the Tinker API.

  • Method summaries at published fidelity

    Summarizing contrastive neuron attribution, Lighthouse Attention, and token superposition using only the stated results.

  • Reproducibility surface

    Pointing to released weights, eval harnesses, and W&B logs where a post says the full stack is public.

  • Attribution and speculation boundary

    Crediting named authors and refusing to extrapolate benchmark numbers or claims beyond the posts.

06

Mission, community, and hiring navigation

Non-technical but load-bearing surface: the open-source mission framing, community and code channels, and the careers funnel with its explicit application format. Covers routing a visitor to the right destination and reproducing stated process exactly.

Mapped capabilities

3 capabilities

  • Channel routing

    Sending a visitor to GitHub, Hugging Face, Discord, or the blog based on what they actually want.

  • Careers application process

    Reproducing the emailed application requirements and the open-ended 'no role fits' path without inventing steps.

  • Mission framing fidelity

    Representing the open-source, human-rights-oriented mission in the company's own terms without adding policy positions.

Coverage is mapped from Nous Research's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Nous Research test?+

The coverage map is generated from Nous Research's own public product surface (open-source language models and distributed AI training infrastructure): 6 scoring areas — Hermes model line and steerability, Hermes Agent, and Psyche distributed training network, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Nous Research evals scored?+

Every case generated for Nous Research — across Hermes model line and steerability and Hermes Agent and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Nous Research library include?+

The full Nous Research library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Variant and size disambiguation and Reasoning vs chat mode behavior under Hermes model line and steerability); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Nous Research or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Nous Research areas and set them up in a Corsac workspace, where you can run every test case against Nous Research or your own agent with your own data.