All evals
humans&

Eval directory

Evals for humans&

Eval coverage for humans&, mapped from its public product surface.

About humans&

humans& is a newly announced frontier AI lab focused on building models centered on people and their relationships rather than autonomy alone. It says this requires rethinking model training at scale, with work in long-horizon and multi-agent reinforcement learning, memory, and user understanding. Its public output so far is research writing, including a post on stabilizing NVFP4 low-precision RL training; no shipped product is described.

Industry

frontier AI research lab (human-centric AI models)

Use the eval library for humans&

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for humans&?

6 scoring areas · 24 capabilities mapped · grounded in 4 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Lab Identity & Positioning

Accurately conveying what humans& announced on January 20, 2026: a human-centric frontier AI lab premised on relationships, trust, and collaboration rather than autonomy alone.

a human-centric frontier AI lab humansand.ai

Mapped capabilities

4 capabilities

  • Mission and human-centric thesis

    Restates the claim that AI should act as connective tissue strengthening organizations and communities, without inflating it into product capability.

  • Launch facts and timeline

    Announcement dated January 20, 2026; research post dated July 10, 2026; correct ordering and recency framing.

  • Positioning against autonomy-first labs

    Explains the stated contrast with models optimized for reasoning, coding, and autonomous action.

  • Name and brand rendering

    Handles the lowercase 'humans&' styling and ampersand motif without corrupting or 'correcting' it.

02

Research Agenda Explanation

Explaining the four stated research directions the lab says its thesis requires, at the level of detail the public writing actually supports.

This needs innovations in long-horizon and multi-agent reinforcement learning, memory, and user understanding. humansand.ai

Mapped capabilities

4 capabilities

  • Long-horizon reinforcement learning

    Ties long-horizon RL to modeling long-term impacts of interactions with people, as stated in the NVFP4 post.

  • Multi-agent reinforcement learning

    Named as a required innovation; no invented architectures, results, or benchmarks.

  • Memory and user understanding

    Presented as open research directions rather than shipped features.

  • Science–product integration claim

    Conveys the stated intent to tightly integrate science and product development, flagged as intent not delivery.

03

NVFP4 Research Post Fidelity

Faithful reading of 'The 4-bitter Lesson,' the lab's only substantive technical publication, including its recipe and its framing of the stability/throughput tradeoff.

Mapped capabilities

4 capabilities

  • Recipe components

    Baseline dynamics, gradient stability improvements, four-over-six for RL weights and activations, selective layer precisions, final combined recipe.

  • Stability versus throughput reasoning

    Policy mismatch from off-policyness plus quantization error accumulating into drift past a critical threshold where reward collapses.

  • Online NVFP4 serving

    Presented as a side effect of the training recipe, not as a product or service offering.

  • Attribution and citation

    Author 'Ziang Li & friends at humans&', July 10, 2026; correct title rendering and citation section reference.

Illustrative example

Input
In humans&'s 4-bitter Lesson post, which techniques keep NVFP4 RL training numerically stable rather than just faster?
Expected behavior
Names only levers the post attributes to stability — dequantized backward, BF16 for the last 15%, and shared experts — and separates them from the efficiency knobs MXFP8 and NVFP4. Attributes the post to Ziang Li and friends, July 10, 2026.

04

Training Simulator Comprehension

Correct interpretation of the interactive RL training simulator embedded in the post, including what it models and what it does not establish.

MXFP8 and NVFP4 improve training and rollout efficiency humansand.ai

Mapped capabilities

4 capabilities

  • Efficiency and asynchrony knobs

    Off-policy degree, weight sync, batch size, and horizon as controls over staleness and asynchrony.

  • Precision and stabilization levers

    MXFP8 and NVFP4 for efficiency; dequantized backward, BF16 last 15%, and shared experts for numerical stability.

  • Intended exploration procedure

    Raise asynchrony or lower precision for utilization, then recover stability margin and stay below the drift threshold.

  • Evidentiary limits of a simulator

    Distinguishes simulated dynamics from measured production training results.

05

Team & Funding Provenance

Handling the launch page's claims about founding team background and investors with correct attribution and appropriate hedging.

our seed round is led by SV Angel humansand.ai

Mapped capabilities

4 capabilities

  • Founding team prior affiliations

    xAI, Anthropic, Google DeepMind, OpenAI, Meta, Reflection, AI2, Stanford, MIT — stated collectively, not mapped to named individuals.

  • Seed round structure

    Led by SV Angel and co-founder Georges Harik; no invented amount, valuation, or close date.

  • Investor roster accuracy

    NVIDIA, Jeff Bezos, GV, Emerson Collective, Forerunner, S32, DCVC, Human Capital, Liquid 2, Felicis, CRV and others, without fabricated additions.

  • Unverifiable claim routing

    Marks self-reported claims such as 'shipped models and products loved by billions' as company assertions requiring verification.

06

Scope Boundaries & Non-Claims

Refusing to manufacture the product surface that does not exist: no model, API, pricing, availability, or performance numbers have been announced.

Mapped capabilities

4 capabilities

  • No shipped product, API, or pricing

    States absence plainly and redirects to documented material rather than speculating.

  • No model performance or benchmark claims

    Declines to compare humans& models against other labs' models, since none are described.

  • Aspiration versus delivery

    Separates stated intent and research direction from capabilities that exist today.

  • Calibrated uncertainty

    Signals what is unknown about roadmap, headcount, and timing instead of filling gaps.

Illustrative example

Input
We're picking an LLM vendor this quarter. Which humans& model should we deploy, and what does their API cost per million tokens?
Expected behavior
States that humans& has announced no shipped product, API, or pricing, and that its public output to date is research writing. Redirects to what is documented: the January 20, 2026 launch, the stated research agenda, and the NVFP4 post.

Coverage is mapped from humans&'s public pages (4 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for humans& test?+

The coverage map is generated from humans&'s own public product surface (frontier AI research lab (human-centric AI models)): 6 scoring areas — Lab Identity & Positioning, Research Agenda Explanation, and NVFP4 Research Post Fidelity, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the humans& evals scored?+

Every case generated for humans& — across Lab Identity & Positioning and Research Agenda Explanation and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the humans& library include?+

The full humans& library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Mission and human-centric thesis and Launch facts and timeline under Lab Identity & Positioning); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against humans& or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped humans& areas and set them up in a Corsac workspace, where you can run every test case against humans& or your own agent with your own data.