All evals
T

Eval directory

Evals for Tabby

Eval coverage for Tabby, mapped from its public product surface.

About Tabby

Tabby is an open-source AI coding assistant that can be self-hosted or run in the cloud, offering code completion, an answer engine, code browsing, and inline chat. Alongside it, TabbyML ships Pochi, an open-source agentic coding teammate that automates multi-step tasks across a whole project and embeds in VS Code and Slack. Plans range from a free Community tier through a per-seat Team tier to a custom Enterprise tier, plus usage-based Tabby Cloud billing.

Industry

self-hosted AI coding assistant / agent

Use the eval library for Tabby

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Tabby?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Code Completion & Inline Assistance

The in-editor assistant surface: tab completion, inline chat, and the editor integrations that carry them.

Mapped capabilities

4 capabilities

  • Tab completion behavior

    Completion suggestions in the editor, including the marketing claim that Tab Completion is always free with no usage limits.

  • Inline chat

    In-editor conversational assistance listed among Tabby's core AI capabilities.

  • Editor and IDE integration

    VS Code and Language Server Protocol based extension paths, plus Cloud IDE support.

  • Telemetry-aware client behavior

    How IDE/extension clients behave under an enforced telemetry policy, where that policy is available.

02

Answer Engine & Code Context

Retrieval and question-answering over a team's own code and documents, including the code browser.

Mapped capabilities

4 capabilities

  • Answer engine responses

    Answering developer queries against connected repositories and context.

  • Code browser

    Browsing and navigating indexed code as a first-class product surface.

  • Context providers and repository context

    Configured context sources, including private GitHub repository connections and repository-level context for completion.

  • Doc ingestion API

    Ingesting a team's own documentation so it can be retrieved alongside code.

03

Pochi Agentic Teammate

The open-source agent that automates multi-step work across a project and embeds in VS Code and Slack.

Tabby is an open-source AI coding assistant, designed to bring the power of AI to your development workflow www.tabbyml.com

Mapped capabilities

4 capabilities

  • Multi-step task execution

    Planning and carrying out multi-step tasks, including coordinated multi-file changes across a workspace.

  • Workflows, rules, and custom agents

    Repeatable step-by-step workflows, project or global rules enforcing conventions, and agents defined with custom prompts and toolsets.

  • MCP and tool use

    Model Context Protocol server configuration connecting APIs, databases, and tools to agent runs.

  • Surfaces and remote execution

    Invocation from VS Code and Slack, and Remote Pochi cloud instances running asynchronously and in parallel.

04

Deployment & Model Configuration

Getting Tabby running on the user's terms: self-hosted, cloud, or air-gapped, on the hardware and models they choose.

Easily integrates with your existing infrastructure, including Cloud IDEs, with support for consumer-grade GPUs. www.tabbyml.com

Mapped capabilities

4 capabilities

  • Self-hosting and cloud options

    Choosing between Tabby Cloud and self-hosted deployment, with no external DBMS or cloud service dependency.

  • Hardware and accelerator support

    Consumer-grade GPUs and documented backends including AMD ROCm, Vulkan, and Metal.

  • Scaled and restricted topologies

    Replicas behind a reverse proxy, and Docker deployment in air-gapped environments.

  • Model selection

    Connecting any AI provider or running local models, including documented integrations such as Codestral.

05

Access, Admin & Governance

Workspace administration and the enterprise controls that gate it.

Mapped capabilities

4 capabilities

  • Secure access and user management

    Member access within the documented per-tier user counts of 5, 50, and unlimited.

  • Authentication domain and SSO

    Domain-based authentication and Single Sign-On, both listed as Enterprise-tier features.

  • Usage reports and analytics

    Administrative reporting on workspace usage.

  • Telemetry policy enforcement

    Enforcing IDE and extension telemetry policy, listed as an Enterprise-tier control.

Illustrative example

Input
We're 20 engineers on the $19/user Team plan. Can we turn on single sign-on and an authentication domain for the workspace?
Expected behavior
States that SSO and Authentication Domain are Enterprise-tier features not included in Team, confirms 20 seats fit within Team's 50-user ceiling, and points to the Enterprise contact path for a custom annually billed plan.

06

Plans, Billing & Usage Limits

Plan boundaries and the usage-based billing model behind Tabby Cloud and Pochi.

Run Tabby in your way and on your terms with no need for external DBMS or cloud services. www.tabbyml.com

Mapped capabilities

4 capabilities

  • Tier comparison and feature gating

    Community at $0 for up to 5 users, Team at $19/user/month for up to 50, and custom annual Enterprise, with per-tier feature availability.

  • Token-based usage billing

    Pochi billing on the token cost of the LLMs run, with $20 in free monthly credits.

  • Charge timing and budgets

    Automatic billing above $10 of usage versus month-end settlement below it, and user-set budgets in the profile.

  • Free-tier boundaries

    Tab Completion remaining free without usage limits, and the stated possibility of a card request under unusually high usage.

Illustrative example

Input
On Tabby Cloud, does tab completion count against my usage bill, and how does that compare to what Pochi charges me?
Expected behavior
Explains that Tab Completion is always free with no usage limits, while Pochi bills on the token cost of the LLMs run, offset by $20 in free monthly credits. Does not attach a per-token price to tab completion.

Coverage is mapped from Tabby's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Tabby test?+

The coverage map is generated from Tabby's own public product surface (self-hosted AI coding assistant / agent): 6 scoring areas — Code Completion & Inline Assistance, Answer Engine & Code Context, and Pochi Agentic Teammate, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Tabby evals scored?+

Every case generated for Tabby — across Code Completion & Inline Assistance and Answer Engine & Code Context and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Tabby library include?+

The full Tabby library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Tab completion behavior and Inline chat under Code Completion & Inline Assistance); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Tabby or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Tabby areas and set them up in a Corsac workspace, where you can run every test case against Tabby or your own agent with your own data.