01
Hermes model line and steerability
Questions about the Hermes series as released: variants and parameter sizes, hybrid reasoning versus chat modes, extended context claims, and how steerability is described. Covers whether the surface answers model-selection questions accurately and separates released facts from speculation.
“Hermes 4.3 was trained with an extended context length (up to 512K)” nousresearch.com
Mapped capabilities
4 capabilities
Variant and size disambiguation
Distinguishing Hermes 4 405B / 70B / 14B, Hermes-4.3-Seed-36B, DeepHermes-3-8B, and Hermes-3-Llama-3.2-3B by size and stated purpose.
Reasoning vs chat mode behavior
Explaining hybrid-mode reasoning and deep-reasoning/chat mode toggles as documented, without inventing configuration flags.
Context length and local-inference fit
Handling the 512K extended-context claim for Hermes 4.3 and which variants are positioned for local hardware.
Steerability framing
Describing Hermes fine-tunes as highly steerable in the terms the company uses, without overclaiming safety or alignment properties.
Illustrative example
- Input
- I have one 24GB GPU. Which Hermes 4 model should I run, and roughly how large is it?
- Expected behavior
- Names Hermes-4-14B as the small, dense variant positioned for local inference and gives its listed size of 28 (GB per the releases table). Does not recommend the 70B or 405B variants for this hardware, and does not invent quantization or throughput numbers.





