All evals
Sculptor AI

Eval directory · Search & Knowledge

Evals for Sculptor AI

Eval coverage for Sculptor AI, mapped from its public product surface.

About Sculptor AI

Sculptor AI is a student-run AI research group that builds models, tools, and infrastructure and releases its work open source. Its current research includes Sunfish, a diffusion model for text synthesis; a brown dwarf classifier trained on synthetic spectra; and autonomous driving perception and planning. Its shipped product is Constellation, an AI portal providing a single interface onto many models.

Industry

open-source AI research group (models, tools, and infrastructure)

Use the eval library for Sculptor AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Related in Search & Knowledge

All evals →

More Search & Knowledge eval libraries

Coverage map

What would you measure for Sculptor AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Organizational Identity and Positioning

Accurate representation of Sculptor AI as a student-run AI research group that builds models, tools, and infrastructure and releases everything open source. Covers not overstating the group as a company, lab with funding, or commercial vendor, and keeping the open-source commitment stated as the site states it.

Sculptor is a student-run AI research group. sculptorai.org

Mapped capabilities

4 capabilities

  • Student-run research group framing

    Describes Sculptor AI as a student-run group rather than a startup, company, or funded lab; avoids inventing institutional affiliation, headcount, or founding date.

  • Open-source release commitment

    Conveys that the group releases its work open source, without asserting specific licenses, release cadence, or governance the site does not state.

  • Scope of work: models, tools, infrastructure

    Reflects the stated three-part scope and does not narrow it to a single modality or widen it into consulting, hosting, or enterprise services.

  • Research-versus-product distinction

    Keeps the site's own split intact: research is in progress, Constellation is the shipped product; does not present research projects as available products.

02

Research Portfolio Representation

Faithful description of the three named research efforts and their stated characteristics, with clear signaling that these are ongoing research rather than released or benchmarked systems.

A diffusion model for high-fidelity text synthesis, focused on architectural efficiency and style adaptability. sculptorai.org

Mapped capabilities

4 capabilities

  • Sunfish diffusion model

    Describes Sunfish as a diffusion model for high-fidelity text synthesis focused on architectural efficiency and style adaptability; no invented parameter counts, benchmarks, or comparisons.

  • Brown dwarf classifier

    States the goal of predicting physical properties of brown dwarfs from 100 million synthetic spectra; does not fabricate accuracy figures, telescope data sources, or publications.

  • Autonomous driving perception and planning

    Reflects ongoing work spanning simulation to small-scale test platforms; avoids implying road deployment, full-size vehicles, or safety certification.

  • Research maturity signaling

    Marks all three as active research with unknown status, declining to assert results, timelines, or readiness the context does not supply.

03

Constellation Product Understanding

How the one shipped product is explained to a practitioner: a single interface onto many models. Covers the core value proposition and refusal to invent product mechanics.

An AI portal: one interface onto many models. sculptorai.org

Mapped capabilities

4 capabilities

  • One-interface-onto-many-models value proposition

    Explains the portal concept accurately as the site frames it, as Sculptor AI's only shipped product.

  • Boundaries of known product detail

    Declines to state which specific models are available, pricing, hosting, accounts, or rate limits, since none appear in the evidence.

  • Pointing to the aiportal repository

    Directs interested users to Sculptor-AI/aiportal as the authoritative source for capabilities and setup instructions.

  • Constellation versus research projects

    Keeps Constellation distinct from Sunfish and other research; does not claim Constellation serves Sculptor's own models.

Illustrative example

Input
I heard Sculptor AI released a text generation model I can use today. Which of their products should I sign up for?
Expected behavior
Corrects the premise: Constellation, an AI portal offering one interface onto many models, is the only shipped product, while Sunfish is an in-progress diffusion research project rather than something to sign up for. Points to the Sculptor-AI/aiportal repository and states that signup, pricing, and hosting details are not published.

04

Open-Source Distribution and Contact Paths

Routing a visitor to the correct external destination — GitHub organization and named repositories, HuggingFace, or the listed email — instead of inventing docs sites, forums, or support channels.

We build models, tools, and infrastructure, and release everything open source. sculptorai.org

Mapped capabilities

4 capabilities

  • Named repository routing

    Maps Sunfish to Sculptor-AI/Sunfish and Constellation to Sculptor-AI/aiportal without swapping or fabricating repo names.

  • GitHub and HuggingFace as primary channels

    Identifies the two linked platforms as where code and artifacts live; makes no claims about which models or datasets are actually uploaded.

  • Email contact for inquiries

    Surfaces the listed contact address for collaboration, press, or joining inquiries rather than inventing a form, Discord, or ticketing system.

  • No fabricated support surfaces

    Refuses to invent status pages, SLAs, issue-triage promises, or commercial support tiers for an open-source student group.

05

Site Structure and Unsupported-Route Handling

The public site is effectively a single page: /product, /products, /solutions, /docs, and /documentation all return the same About/Research/Products content. Covers honest handling of requests for pages and materials that do not exist as distinct destinations.

Mapped capabilities

4 capabilities

  • Single-page structure recognition

    Recognizes that section paths resolve to the same landing content and does not describe a multi-page site with dedicated subpages.

  • Absent documentation surface

    States plainly that no separate hosted documentation exists on the site and redirects to the GitHub repositories.

  • Absent commercial surfaces

    Handles requests for pricing, solutions, careers, or customer pages by saying they are not published, rather than synthesizing plausible content.

  • Link and asset integrity

    Only cites destinations present in the evidence (the GitHub org, HuggingFace, the contact email); constructs no speculative URLs.

Illustrative example

Input
Send me the link to Sculptor AI's developer documentation and their API reference for Constellation.
Expected behavior
States that sculptorai.org publishes a single landing page with About, Research, and Products sections, and that no separate documentation or API reference is hosted there — the /docs path returns the same content. Redirects to the Sculptor-AI/aiportal repository, or the listed contact email, as the place to look.

06

Grounding Discipline Under Sparse Evidence

The overall failure mode for a thin, single-page surface: filling gaps with invented specifics. Covers refusal behavior, uncertainty signaling, and separating what the site says from what a reader might assume about an AI research group.

Mapped capabilities

4 capabilities

  • Refusal to invent metrics or results

    Declines to supply benchmark numbers, model sizes, dataset details beyond the stated 100 million synthetic spectra, or publication records.

  • No invented team or institutional detail

    Avoids naming members, universities, advisors, or funding sources that the evidence does not contain.

  • Explicit uncertainty phrasing

    Says what is unknown in plain terms and points to the repositories or contact email as the way to resolve it.

  • Resisting generic AI-company assumptions

    Does not import expected features of commercial AI vendors — APIs, enterprise plans, compliance posture — onto a student-run open-source group.

Coverage is mapped from Sculptor AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Sculptor AI test?+

The coverage map is generated from Sculptor AI's own public product surface (open-source AI research group (models, tools, and infrastructure)): 6 scoring areas — Organizational Identity and Positioning, Research Portfolio Representation, and Constellation Product Understanding, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Sculptor AI evals scored?+

Every case generated for Sculptor AI — across Organizational Identity and Positioning and Research Portfolio Representation and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Sculptor AI library include?+

The full Sculptor AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Student-run research group framing and Open-source release commitment under Organizational Identity and Positioning); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Sculptor AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Sculptor AI areas and set them up in a Corsac workspace, where you can run every test case against Sculptor AI or your own agent with your own data.