All evals
A

Eval directory

Evals for Adaption

Eval coverage for Adaption, mapped from its public product surface.

About Adaption

Adaption (Adaption Labs) is building adaptability-first AI systems that continually learn instead of relying on large, static, one-size-fits-all models. The company positions itself against brute-force scaling, emphasizing dynamically shaped data, gradient-free and continual learning. The public site is largely a manifesto plus a careers page and a login/contact entry point; no specific product feature set, pricing, or compliance details are published.

Industry

adaptive AI / continual-learning foundation models

Use the eval library for Adaption

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Adaption?

5 scoring areas · 19 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Positioning & Manifesto Fidelity

Faithfully explaining Adaption's stated thesis — adaptability-first systems that continually learn, an explicit bet against brute-force scaling, and the critique of monolithic one-size-fits-all models — without embellishing it into product claims.

We are betting against scaling, and instead building efficient AI that continually learns. adaptionlabs.ai

Mapped capabilities

4 capabilities

  • Core thesis restatement

    Summarize the 'betting against scaling' argument and the case that averages erase exceptional or edge use cases.

  • Technical vocabulary handling

    Explain the site's named concepts — dynamically shaped data, gradient-free learning, continual learning — at the level of specificity the site actually provides.

  • Problem framing

    Convey the stated user pain (prompt engineering, contorting requests, slow expensive update cycles) in the company's own framing.

  • Claim boundary

    Present manifesto statements as company positioning rather than as demonstrated results or benchmarks.

02

Careers & Role Discovery

Accurate retrieval and filtering over the published careers listing: role titles, owning teams, locations, and employment or work-arrangement types.

Mapped capabilities

4 capabilities

  • Role and team enumeration

    Report the 12 open roles and the teams they belong to (Modelling, Platform, GTM, Applied ML, Marketing, Operations).

  • Location filtering

    Answer location-scoped questions such as which roles are open in San Francisco, Singapore, or globally remote.

  • Employment type and arrangement

    Distinguish FullTime from Contract roles and Remote from Hybrid and OnSite arrangements, including named contract durations.

  • Application routing

    Direct a candidate to the per-role Apply path on the careers page rather than to a generic contact form.

Illustrative example

Input
Which of Adaption's open roles are contract positions rather than full-time, and where is each one based?
Expected behavior
Identifies the three contract roles — Modelling Resident (8-month), Content Marketing Manager (6-month), and Recruiter (8-month) — with their listed locations, and does not misclassify the full-time Growth Generalist despite its stated 8-month framing.

03

Access, Login & Inbound Routing

Correctly routing visitors among the gated app entry point, the login affordance present across the site, and the contact page — recognizing that the product itself is behind authentication.

Mapped capabilities

4 capabilities

  • Login vs. app entry

    Point to the app authentication surface for existing users and note that the product is not publicly browsable.

  • Contact and learn-more routing

    Route prospective customers and general inquiries to the Contact Us / learn-more page.

  • No self-serve signup claims

    Avoid asserting a public signup, free trial, or waitlist flow that the site does not document.

  • Intent disambiguation

    Separate candidate intent, customer intent, and existing-user access intent onto the right destination.

05

Unpublished-Detail Discipline

Handling the large set of questions the public surface simply does not answer — pricing, feature lists, model specs, customers, security certifications, funding — with explicit non-availability instead of plausible invention.

Mapped capabilities

4 capabilities

  • Pricing and packaging

    State that no pricing, tiers, or commercial terms are published and route to contact.

  • Product feature and spec claims

    Avoid describing capabilities, APIs, integrations, or model specifications that the site never lists.

  • Customers, benchmarks, and metrics

    Decline to name customers, cite performance numbers, or quantify efficiency claims the site states only qualitatively.

  • Company facts beyond the site

    Separate what the public surface supports from unverified external claims about funding, headcount, or leadership.

Illustrative example

Input
What does Adaption charge for its adaptive AI platform, and is it SOC 2 certified?
Expected behavior
States plainly that Adaption publishes no pricing, plans, or compliance certifications on its public site, offers no estimate or guess, and points the user to the Contact Us page as the available route for commercial and security questions.

Coverage is mapped from Adaption's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Adaption test?+

The coverage map is generated from Adaption's own public product surface (adaptive AI / continual-learning foundation models): 5 scoring areas — Positioning & Manifesto Fidelity, Careers & Role Discovery, and Access, Login & Inbound Routing, and more — spanning 19 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Adaption evals scored?+

Every case generated for Adaption — across Positioning & Manifesto Fidelity and Careers & Role Discovery and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Adaption library include?+

The full Adaption library is built on request. The coverage map spans 5 areas and 19 capabilities (for example, Core thesis restatement and Technical vocabulary handling under Positioning & Manifesto Fidelity); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Adaption or my own agent?+

Request the library with your work email above. We'll build out all 5 mapped Adaption areas and set them up in a Corsac workspace, where you can run every test case against Adaption or your own agent with your own data.