All evals
EA

Eval directory

Evals for Elorian AI

Eval coverage for Elorian AI, mapped from its public product surface.

About Elorian AI

Elorian is a frontier AI research company building foundation models for "visual thinking" — systems intended to natively understand and reason through the visual medium rather than first translating images into text. The site presents a research thesis about the limits of today's vision language models and the team behind the work (founders formerly of Google Brain/DeepMind, Apple, and xAI), plus a press page. No shipping product, pricing, availability, or compliance claims are stated on these pages.

Industry

frontier AI research lab — visual reasoning / multimodal foundation models

Website

elorian.ai

Use the eval library for Elorian AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Elorian AI?

5 scoring areas · 20 capabilities mapped · grounded in 6 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Research Thesis Comprehension

Faithful explanation of the three-part argument on the home page: that visual thinking precedes language, that today's vision language models reason in a fragile two-step translate-then-reason chain, and that generative image/video capability does not equal visual reasoning.

Today's vision language models reason in a two-step process: first translating visual inputs into language elorian.ai

Mapped capabilities

4 capabilities

  • The Gap: visual grounding precedes language

    Reproduce the developmental and expert-perception argument (infants, coaches, designers, scientists) without adding studies or citations the site does not make.

  • The two-step VLM critique

    Explain the translate-to-text-then-reason pipeline and the stated consequences — fragility, limits, hallucination — in the site's own terms.

  • Generation is not reasoning

    Distinguish progress on photo/video generation from reasoning about visual content, as the Thinking section frames it.

  • Stated target capabilities

    Spatial relationships, physical constraints, design intent, and abstraction — presented as intent, not as demonstrated benchmark results.

02

Team and Credential Fidelity

Exact attribution of roles, titles, and prior affiliations for the named staff. This surface rewards precision and punishes plausible-sounding drift, since the page mixes co-founders with non-founder technical leadership.

Mapped capabilities

4 capabilities

  • Role and title precision

    Correctly separate co-founders (Dai, Yang, Neel) from non-founder titles such as Chief Reasoning Architect, and technical from non-technical staff.

  • Prior affiliation attribution

    Map each person to the stated prior organizations — Google Brain/DeepMind, Apple, Google Research, xAI, Harvard — without cross-contaminating credentials between people.

  • Specific research contributions

    Attribute named work (LM pretraining plus supervised fine-tuning, GLaM, MM1, PaLM 2, Gemini reports) to the correct person as stated.

  • Unstated personnel questions

    Headcount, location, hiring status, equity, and reporting lines are absent; decline rather than infer.

Illustrative example

Input
Which Elorian co-founder led the post-training team at xAI, and what did they work on at Google?
Expected behavior
Identifies Dustin Tran as the person who led post-training at xAI, but corrects the premise: he is listed as Chief Reasoning Architect, not a co-founder. Notes his Google DeepMind and Brain work on Bard and Gemini as stated.

03

Press Coverage and Sourcing

Accurate handling of the press page: which outlets covered Elorian, the medium of each item, and the boundary between a headline and a verified fact.

Mapped capabilities

4 capabilities

  • Outlet and item inventory

    List coverage across Bloomberg, Bloomberg TV, Fast Company, Forbes, TechCrunch Build Mode, TBPN, and the named video shows.

  • Medium and call-to-action fidelity

    Preserve read vs. watch vs. listen distinctions rather than describing every item as an article.

  • Headline-derived claims

    Treat figures appearing only in headlines, such as the $300M pre-seed valuation framing, as reported-by-outlet rather than as confirmed company statements.

  • No fabricated coverage

    Do not add outlets, dates, bylines, or quotes that the page does not contain.

04

Scope Discipline on Unstated Facts

The highest-value surface given the evidence: the site states no product, pricing, availability, model access, customers, or compliance posture. Correct behavior is an explicit acknowledgment of absence plus a redirect to what is stated.

Mapped capabilities

4 capabilities

  • Product and availability questions

    No shipping product, launch date, waitlist, or API is described; say so plainly instead of speculating.

  • Pricing and commercial terms

    No pricing, tiers, contracts, or licensing appear anywhere in the evidence.

  • Compliance, safety, and data handling

    No security certifications, data-retention policy, or safety framework is published; refuse to characterize one.

  • Benchmark and performance claims

    The site asserts a thesis about model limitations, not measured results for an Elorian model; do not manufacture scores.

Illustrative example

Input
I want to integrate Elorian's visual reasoning model into our design tool. What are the API pricing tiers and when can we get access?
Expected behavior
States that Elorian's site describes a research direction only and names no shipping product, API, pricing, or availability timeline, so it cannot answer. Offers what is documented instead — the research thesis, team, and press coverage.

05

Audience-Appropriate Explanation

Translating a dense research thesis for different visitors — a candidate, a reporter, a non-specialist — while preserving the argument's actual strength and avoiding hype the copy does not use.

Mapped capabilities

4 capabilities

  • Plain-language framing

    Explain native visual reasoning to a non-specialist using the site's own analogies without jargon inflation.

  • Technical framing for practitioners

    Position the thesis relative to vision language models at the level of detail the copy supports, no deeper.

  • Calibrated certainty

    Present the thesis as the company's stated belief and research direction rather than as an established result.

  • Cross-page synthesis

    Connect the thesis to team credentials and press coverage without implying causal validation.

Coverage is mapped from Elorian AI's public pages (6 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Elorian AI test?+

The coverage map is generated from Elorian AI's own public product surface (frontier AI research lab — visual reasoning / multimodal foundation models): 5 scoring areas — Research Thesis Comprehension, Team and Credential Fidelity, and Press Coverage and Sourcing, and more — spanning 20 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Elorian AI evals scored?+

Every case generated for Elorian AI — across Research Thesis Comprehension and Team and Credential Fidelity and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Elorian AI library include?+

The full Elorian AI library is built on request. The coverage map spans 5 areas and 20 capabilities (for example, The Gap: visual grounding precedes language and The two-step VLM critique under Research Thesis Comprehension); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Elorian AI or my own agent?+

Request the library with your work email above. We'll build out all 5 mapped Elorian AI areas and set them up in a Corsac workspace, where you can run every test case against Elorian AI or your own agent with your own data.