All evals
Plastic Labs

Eval directory

Evals for Plastic Labs

Eval coverage for Plastic Labs, mapped from its public product surface.

About Plastic Labs

Honcho is Plastic Labs' AI-native memory product, described as a continual learning system for modeling personal identity and, in the future, a shared context layer for individual alignment. It is powered by the company's Neuromancer model series, reasoning models built for memory and social cognition. Neuromancer XR, a fine-tuned Qwen3-8B model for explicit reasoning, is currently deployed in Honcho, while a meta-reasoning model (MR) is in training.

Industry

AI-native memory layer for LLM agents

Use the eval library for Plastic Labs

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Plastic Labs?

6 scoring areas · 21 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Identity Representation & Social Reasoning

Forming representations of personal identity by reasoning over user and agent data, the core function Honcho is described as performing.

Honcho is a continual learning system for modeling personal identity plasticlabs.ai

Mapped capabilities

4 capabilities

  • Deriving traits and facts from dialogue evidence

    Attributes that follow from what the user actually said, not from stereotype or priors.

  • Separating derived conclusions from speculation

    Deductions stated as such; unsupported guesses withheld or labeled.

  • Evidence traceability

    Each represented fact points back to the source turns that support it.

  • Entity scope: user versus agent representations

    Representations are described as covering any entity; keeping them distinct.

Illustrative example

Input
Stored turns: "Our team ships every Tuesday and I run the release." and "I've never worked a job outside backend." Question: what does this user own?
Expected behavior
States that the user owns the weekly Tuesday release as a backend engineer, and cites both stored turns as the support. It adds no seniority, employer, or tenure detail that the two turns do not entail.

02

Continual Learning & Memory Maintenance

Honcho is described as a continual learning system, so representations must change correctly as new information arrives rather than freezing at first observation.

an AI-native memory solution powered by our state-of-the-art reasoning models plasticlabs.ai

Mapped capabilities

4 capabilities

  • Contradiction and supersession handling

    Newer information updates the current view without silently erasing history.

  • Durable traits versus transient state

    Stable identity signals are not overwritten by one-off remarks.

  • Retention across sessions

    Prior context remains available and correctly attributed later.

  • Unknowns stay unknown

    No fabricated attribute when the stored evidence does not cover the question.

Illustrative example

Input
Stored earlier: "I'm a Python backend engineer." Stored later: "I moved into engineering management last month." Question: what is this user's current role?
Expected behavior
Reports engineering management as the current role and keeps the earlier backend engineer fact as superseded prior history rather than deleting it or presenting it as current.

03

Explicit Reasoning Quality (Neuromancer XR)

XR is the currently deployed model, a fine-tuned Qwen3-8B trained on social reasoning traces and described as excelling at deduction and factual derivation toward certain conclusions.

NEUROMANCER is the first collection of models dedicated to AI-native memory and social cognition. plasticlabs.ai

Mapped capabilities

4 capabilities

  • Formal logical steps over user and agent data

    Multi-step deduction that holds together, not restatement of inputs.

  • Certain versus uncertain conclusions

    Committing where the evidence entails it, hedging where it does not.

  • Legibility of the reasoning trace

    A reviewer can follow how a conclusion was reached.

  • Stability across repeated queries

    The same stored evidence yields the same derived conclusion.

04

Predictive Reasoning Boundaries (Neuromancer MR)

MR, the meta-reasoning model for inductions and abductions over XR traces, is stated to be in training. This area covers how the product behaves at that boundary today.

trained specifically to do formal logical reasoning over user and agent data plasticlabs.ai

Mapped capabilities

3 capabilities

  • Hypotheses labeled as hypotheses

    Pattern-level or predictive claims are not presented as derived fact.

  • No implied MR-only behavior

    Capability described as in training is not claimed as deployed.

  • Escalation from trace to pattern claim

    Where a run moves from deduction to conjecture is visible.

05

Grounded Q&A over Published Content

Plastic's blog exposes an 'Ask Honcho' chat over its posts, research, and notes, plus markdown and copy-all views of the same corpus.

Mapped capabilities

4 capabilities

  • Answers restricted to the published corpus

    Questions about the blog answered from the blog, not from general recall.

  • Source attribution

    Answers point to the post, research piece, or note relied on.

  • Deferral outside the corpus

    Off-corpus questions get a clear 'not covered here' rather than an invention.

  • Accuracy on stated product and model facts

    XR's base model, its deployment status, and MR's training status reported correctly.

06

Shared Context Layer (Forthcoming)

Described as a near-term direction rather than a shipped capability: a shared context layer for individual alignment across applications. Scoped here as roadmap coverage only.

soon a shared context layer for individual alignment plasticlabs.ai

Mapped capabilities

2 capabilities

  • Roadmap claim discipline

    Forthcoming capability described as forthcoming when asked.

  • Per-entity context boundaries

    One entity's representation does not leak into another's.

Coverage is mapped from Plastic Labs's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Plastic Labs test?+

The coverage map is generated from Plastic Labs's own public product surface (AI-native memory layer for LLM agents): 6 scoring areas — Identity Representation & Social Reasoning, Continual Learning & Memory Maintenance, and Explicit Reasoning Quality (Neuromancer XR), and more — spanning 21 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Plastic Labs evals scored?+

Every case generated for Plastic Labs — across Identity Representation & Social Reasoning and Continual Learning & Memory Maintenance and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Plastic Labs library include?+

The full Plastic Labs library is built on request. The coverage map spans 6 areas and 21 capabilities (for example, Deriving traits and facts from dialogue evidence and Separating derived conclusions from speculation under Identity Representation & Social Reasoning); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Plastic Labs or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Plastic Labs areas and set them up in a Corsac workspace, where you can run every test case against Plastic Labs or your own agent with your own data.