All evals
Goodfire

Eval directory

Evals for Goodfire

Eval coverage for Goodfire, mapped from its public product surface.

About Goodfire

Silico is Goodfire's platform for AI research, letting teams inspect what their models have learned, diagnose failures, and make targeted interventions. It pairs a research agent that plans and runs long-horizon experiments with interpretability and training libraries plus compute orchestration. It is offered for large language models, life sciences models, and robotics/vision models.

Industry

AI interpretability and model research platform

Headquarters

San Francisco (Telegraph Hill)

Use the eval library for Goodfire

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Goodfire?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Research Agent & Long-Horizon Experiments

The agent's ability to take a research goal, develop an experimental plan, execute work in parallel, monitor progress, and return inspectable results the researcher can build on.

Give Silico a research goal. It develops an experimental plan, runs the work in parallel www.goodfire.com

Mapped capabilities

4 capabilities

  • Goal-to-plan decomposition

    Turns a stated research goal into an explicit experimental plan before execution begins.

  • Parallel execution and progress monitoring

    Runs planned work concurrently and reports run status without constant supervision.

  • Inspectable, buildable results

    Returns results a researcher can inspect, extend, and iterate on rather than opaque summaries.

  • Paper replication and extension

    Given a paper to reproduce or modify, plans and runs the experiments and compares against the original.

Illustrative example

Input
Reproduce the probe-as-reward hallucination result on a smaller open model, then tell me whether standard performance benchmarks degraded.
Expected behavior
Silico returns an experimental plan naming the base model, probe setup, and benchmarks before launching runs, then reports the hallucination-rate change alongside benchmark scores for both the baseline and trained checkpoint.

02

Interpretability & Model Understanding

Tooling to surface what a model has learned: architecture visualization, SAE and probe training, neural geometry mapping, and causal hypothesis testing.

Mapped capabilities

4 capabilities

  • SAE and probe training

    Trains sparse autoencoders and probes on model internals via the built-in interpretability libraries.

  • Architecture and neural geometry visualization

    Visualizes model structure and maps internal representational geometry.

  • Causal hypothesis testing

    Tests causal claims about internal mechanisms rather than reporting correlational feature descriptions.

  • Feature interpretation quality

    Produces human-meaningful accounts of discovered features tied to observed model behavior.

03

Failure Diagnosis & Pre-Deployment Anticipation

Tracing regressions and unexpected behavior back to concrete causes, and surfacing failure modes that standard benchmarks and eval pipelines miss.

Mapped capabilities

4 capabilities

  • Regression tracing

    Traces an observed behavior change to undertraining, bottlenecks, feature collapse, or spurious correlations.

  • Dataset artifact attribution

    Identifies training datapoints responsible for a specific undesired behavior.

  • Pre-deployment failure prediction

    Surfaces likely failure modes ahead of deployment rather than after incidents.

  • Post-training side-effect detection

    Flags undesired behaviors that emerge during post-training and are hard to detect otherwise.

04

Training & Targeted Intervention

Running SFT, DPO, and RL experiments, comparing checkpoints, applying targeted interventions, and measuring what actually changed.

Mapped capabilities

4 capabilities

  • SFT / DPO / RL experiment runs

    Configures and executes supervised, preference, and reinforcement post-training experiments.

  • Checkpoint comparison and change measurement

    Compares checkpoints and reports what measurably changed after an intervention.

  • Correction without full retraining

    Applies targeted fixes to identified failure modes instead of retraining from scratch.

  • Representation-based reward signals

    Uses internal representations or probes as training reward signals.

05

Compute Orchestration & Environment

Coordinating experiments across compute fabrics, keeping long-running jobs moving, and supporting either Goodfire infrastructure or a customer-connected cluster.

Run on Goodfire's infrastructure or connect your own cluster www.goodfire.com

Mapped capabilities

4 capabilities

  • Cross-fabric experiment coordination

    Schedules and coordinates experiments across compute fabrics.

  • Unattended long-running job handling

    Monitors runs and keeps long-horizon work progressing without babysitting.

  • Bring-your-own-cluster setup

    Runs on Goodfire infrastructure or connects to a customer's own cluster.

  • Frontier-scale activation harvesting

    Handles interpretability data collection at very large model scale.

06

Access, Plans & Data Governance

Plan boundaries, usage and seat handling, and the security and data-retention commitments that gate team adoption.

Zero Data Retention available www.goodfire.com

Mapped capabilities

4 capabilities

  • Individual vs Enterprise boundaries

    Correctly distinguishes single-researcher access from pooled usage, org billing, and seat management.

  • Weekly usage refresh handling

    Explains how usage refreshes and what happens at plan limits.

  • Zero Data Retention and SOC 2 posture

    Represents ZDR availability and SOC 2 Type II certification accurately and without overstatement.

  • Model-type coverage claims

    Correctly scopes support across LLM, life sciences, and robotics/vision models.

Illustrative example

Input
I'm on the Individual plan at $1,000 a month. Can I add two teammates, share my weekly usage with them, and get one org-level invoice?
Expected behavior
Silico states that Individual covers a single researcher with weekly-refreshed usage, and that pooled team usage, org-level billing, and seat management are Enterprise features on custom pricing, then points to requesting a demo.

Coverage is mapped from Goodfire's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Goodfire test?+

The coverage map is generated from Goodfire's own public product surface (AI interpretability and model research platform): 6 scoring areas — Research Agent & Long-Horizon Experiments, Interpretability & Model Understanding, and Failure Diagnosis & Pre-Deployment Anticipation, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Goodfire evals scored?+

Every case generated for Goodfire — across Research Agent & Long-Horizon Experiments and Interpretability & Model Understanding and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Goodfire library include?+

The full Goodfire library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Goal-to-plan decomposition and Parallel execution and progress monitoring under Research Agent & Long-Horizon Experiments); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Goodfire or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Goodfire areas and set them up in a Corsac workspace, where you can run every test case against Goodfire or your own agent with your own data.