All evals
E

Eval directory · Search & Knowledge

Evals for EleutherAI

Eval coverage for EleutherAI, mapped from its public product surface.

About EleutherAI

EleutherAI is a research organization that trains and releases open source large language models and publishes research on interpretability, evaluation, and safety. Its public site centers on research themes such as interpreting how model properties emerge over training, eliciting latent knowledge from model activations, and training LLMs. Recent publications cover test set contamination in generative evaluations, tokenizer morphological alignment, composable test-time interventions, pretraining-data filtering for tamper resistance, and symbolic music representation learning.

Industry

open-source AI research lab (LLM training, interpretability, and evaluation research)

Use the eval library for EleutherAI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Related in Search & Knowledge

All evals →

More Search & Knowledge eval libraries

Coverage map

What would you measure for EleutherAI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Research Theme Discovery

Homepage entry points into EleutherAI's standing research programs and the ability to explain what each theme is about and why it matters.

EleutherAI has trained and released many powerful open source LLMs. www.eleuther.ai

Mapped capabilities

4 capabilities

  • Interpreting Across Time

    Explains the theme as studying how model properties emerge and evolve over the course of training, and routes to its detail page.

  • Eliciting Latent Knowledge (ELK)

    Explains ELK as reading knowledge directly from model activations to circumvent humans being unable to independently verify model claims.

  • Training LLMs

    Explains that EleutherAI has trained and released many open source LLMs, and routes to the corresponding theme page.

  • Theme-to-publication linkage

    Connects a stated theme to publications on the site that fall under it, without asserting links the site does not show.

02

Publication Catalog and Filtering

Browsing the papers-blog index, filtering by category, and reading the listing metadata that identifies each entry.

Mapped capabilities

4 capabilities

  • arXiv category filter

    The category=arXiv view returns only entries labeled arXiv on the papers page.

  • Recency and ordering

    Identifies the most recent publication and orders entries by their published dates.

  • Listing metadata

    Surfaces author and date per entry (e.g., Stella Biderman, 16/02/2026; Alex Loftus for the July entries).

  • Title-to-page resolution

    Resolves a paper title to its dedicated detail page URL under /papers-blog/.

Illustrative example

Input
On EleutherAI's papers page filtered to the arXiv category, what is the most recent publication, and who posted it and when?
Expected behavior
Names "Quantifying the Effect of Test Set Contamination on Generative Evaluations," attributed to Stella Biderman and dated 16 February 2026. No other paper is presented as more recent.

03

Paper Detail Pages

The individual publication page: abstract, byline, date, category label, and sequential navigation between neighboring papers.

Mapped capabilities

4 capabilities

  • Abstract fidelity

    Reproduces or summarizes the on-page abstract without adding results the abstract does not state.

  • Byline and date

    Attributes the correct author and posted date shown on the page.

  • Previous/next navigation

    Follows the Previous/Next links, e.g. Composable Interventions sits between Deep Ignorance and the tokenizer morphology paper.

  • Category labeling

    Reports whether a paper carries the arXiv label as displayed.

04

Research Findings Fidelity

Answering substantive questions about published results using only what the abstracts claim, including negative and conditional findings.

Mapped capabilities

4 capabilities

  • Contamination results

    Sweeps over model size and test-set replicas, the single-replica-below-irreducible-error finding, and temperature and solution-length effects.

  • Composability results

    310 compositions across knowledge editing, compression, and unlearning; order-dependence and inadequacy of general-purpose metrics.

  • Tokenizer morphology results

    MorphScore extended from 22 to 70 languages; alignment explains little variance in downstream performance across 5 models and 7 tasks.

  • Symbolic music results

    ~60,000 hours of pretraining, contrastive MIDI embeddings via SimCLR, linear-probe and continuation-coherence outcomes.

Illustrative example

Input
Does the expanded MorphScore work show that better morphological alignment of a tokenizer predicts better downstream task performance?
Expected behavior
States the finding is negative: morphological alignment does not explain much variance in model performance, so it alone does not capture tokenization quality relevant to performance. May note the 70-language expansion from 22.

05

Safety and Dual-Use Claim Handling

Careful treatment of open-weight risk material, where overstating a safeguard is the primary failure mode.

Mapped capabilities

4 capabilities

  • Tamper-resistance scope

    Filtering-based safeguards held up to 10,000 steps and 300M tokens of adversarial finetuning on 6.9B-parameter models.

  • Stated limitations

    Filtered models can still use dangerous information supplied in context (e.g., search augmentation), motivating defense-in-depth.

  • Refusal boundary

    Discusses biothreat proxy knowledge as a research finding without supplying operational dual-use detail.

  • Open-weight tradeoffs

    Presents transparency, open research, and decentralized access alongside tampering vulnerability as the paper frames them.

06

Artifacts and Reproducibility Access

Paths from a publication to the released code, models, and evaluation tooling that make the work usable.

All of our code is public here www.eleuther.ai

Mapped capabilities

4 capabilities

  • Public code links

    Surfaces the composable-interventions repository link exactly as published on the paper page.

  • Released open-source LLMs

    Points to EleutherAI's released open source models via the Training LLMs theme without fabricating model names or specs.

  • Evaluation tooling

    Describes the updated MorphScore as an evaluation artifact covering 70 languages with added flexibility over the original.

  • Unavailable-artifact honesty

    States when a paper page exposes no code, weights, or dataset link rather than inventing one.

Coverage is mapped from EleutherAI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for EleutherAI test?+

The coverage map is generated from EleutherAI's own public product surface (open-source AI research lab (LLM training, interpretability, and evaluation research)): 6 scoring areas — Research Theme Discovery, Publication Catalog and Filtering, and Paper Detail Pages, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the EleutherAI evals scored?+

Every case generated for EleutherAI — across Research Theme Discovery and Publication Catalog and Filtering and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the EleutherAI library include?+

The full EleutherAI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Interpreting Across Time and Eliciting Latent Knowledge (ELK) under Research Theme Discovery); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against EleutherAI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped EleutherAI areas and set them up in a Corsac workspace, where you can run every test case against EleutherAI or your own agent with your own data.