All evals
LM Studio

Eval directory

Evals for LM Studio

Eval coverage for LM Studio, mapped from its public product surface.

About LM Studio

LM Studio Bionic is a standalone desktop AI agent built around open models, separate from the main LM Studio app. It supports Work Projects for research, writing, and document tasks and Code Projects with file, search, Git, and shell tools, plus real-time local voice transcription. Models can run locally via the LM Studio runtime (MLX and llama.cpp), remotely through LM Link, or on LM Studio's US-based cloud with zero data retention and pay-as-you-go token pricing.

Industry

local AI agent for open models (desktop app)

Use the eval library for LM Studio

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for LM Studio?

6 scoring areas · 22 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Projects and sessions

How Bionic organizes work into Work Projects and Code Projects, with separate sessions per task and multiple sessions running in parallel.

Code Projects for working in a local codebase with file, search, Git, and shell tools. lmstudio.ai

Mapped capabilities

4 capabilities

  • Choosing Work vs Code project type

    Routing a stated task to the right project type: research, writing, analysis, and document work vs working in a local codebase.

  • Creating a first project

    Walking a new user through project creation as described in the Bionic docs, without inventing setup steps.

  • Session scoping

    Keeping separate tasks in separate sessions within a project instead of merging unrelated work.

  • Parallel sessions across projects

    Explaining how multiple sessions run at once to parallelize workflows.

02

Document creation and editing

Work Project behavior for creating and editing documents, including the automatic save model that lets users work with the agent freely.

Mapped capabilities

3 capabilities

  • Create and edit documents in a session

    Producing and revising document content as part of a Work Project task.

  • Automatic saving of changes

    Correctly describing that every change is saved automatically rather than requiring manual save.

  • Research to deliverable

    Turning research material into a finished written artifact within a Work Project.

03

Code Project tooling

Agentic work in a local codebase using the file, search, Git, and shell tools Bionic exposes, plus automations and computer control.

Bionic is LM Studio's agent for open models. Natively local. lmstudio.ai

Mapped capabilities

4 capabilities

  • File and search tools

    Locating and reading the right files in a local codebase before making changes.

  • Git operations

    Using the Git tool for repository state and history as part of a coding task.

  • Shell tool use

    Running commands to complete a coding or automation task, and reporting what was run.

  • Automations and computer control

    Handling requests to automate a local task or control the computer as described on the product page.

04

Model placement: local, remote, and cloud

Choosing where a session's model runs — local models on the LM Studio runtime (MLX, llama.cpp), remote models via LM Link, or frontier open models in LM Studio Cloud.

LM Studio Cloud inference is US-based, with Zero Data Retention (ZDR) by default. lmstudio.ai

Mapped capabilities

4 capabilities

  • Selecting cloud, local, or remote for a session

    Matching model placement to a stated constraint such as offline work, device capability, or task difficulty.

  • Downloading local models in-app

    Guiding a user to download and use the latest local LLMs directly within Bionic.

  • Frontier open models in cloud

    Naming the cloud-available open models the site lists and when heavier tasks warrant them.

  • LM Link remote devices

    Explaining remote model use from another device, including the five-device limit stated on the pricing page.

Illustrative example

Input
I'm drafting a confidential contract summary and nothing can leave my laptop. Set me up in Bionic, and I'd like to dictate the notes instead of typing.
Expected behavior
Recommends a Work Project using a local model on the LM Studio runtime, and confirms voice input is transcribed locally so audio never leaves the device. Flags that cloud models and the web search tool involve sending data off-device, even under Zero Data Retention.

05

Voice input and transcription

Real-time local speech transcription in Bionic, including its multi-language support and on-device processing guarantee.

Your voice and audio data is processed locally and never leaves your device. lmstudio.ai

Mapped capabilities

3 capabilities

  • Real-time transcription during a session

    Transcribing speech as the user talks and acting on it in the session.

  • On-device voice processing

    Stating that voice and audio data are processed locally and never leave the device.

  • Multi-language speech

    Handling supported languages other than English in voice input.

06

Privacy, plans, and spend

Data-handling commitments and commercial boundaries: Zero Data Retention across cloud services, US-based inference, the free local tier, and pay-as-you-go token pricing with credits.

Our cloud services are Zero Data Retention (ZDR) across the board. lmstudio.ai

Mapped capabilities

4 capabilities

  • Zero Data Retention claims

    Describing ZDR accurately for cloud services and the logged-in web search tool, without overstating it.

  • Free tier boundaries

    Distinguishing what runs free and local from what consumes cloud credits.

  • Token pricing and credits

    Reporting per-million-token input, cached, and output prices for listed cloud models and how credits are added.

  • Unreleased plan details

    Declining to invent details about Bionic Pass, which the pricing page marks as coming soon.

Illustrative example

Input
I'll be sending about 2 million input tokens a month to GLM-5.2 in Bionic. What will that cost, and can I just get it on Bionic Pass instead?
Expected behavior
Quotes GLM-5.2 input at $1.50 per million tokens, giving roughly $3.00 for 2M uncached input tokens, and notes output is billed separately at $4.50. States that Bionic Pass pricing and plan details are not yet announced rather than estimating them.

Coverage is mapped from LM Studio's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for LM Studio test?+

The coverage map is generated from LM Studio's own public product surface (local AI agent for open models (desktop app)): 6 scoring areas — Projects and sessions, Document creation and editing, and Code Project tooling, and more — spanning 22 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the LM Studio evals scored?+

Every case generated for LM Studio — across Projects and sessions and Document creation and editing and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the LM Studio library include?+

The full LM Studio library is built on request. The coverage map spans 6 areas and 22 capabilities (for example, Choosing Work vs Code project type and Creating a first project under Projects and sessions); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against LM Studio or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped LM Studio areas and set them up in a Corsac workspace, where you can run every test case against LM Studio or your own agent with your own data.