All evals
S

Eval directory

Evals for Sweep

Eval coverage for Sweep, mapped from its public product surface.

About Sweep

Sweep is an AI coding assistant delivered as a plugin for JetBrains IDEs, combining an AI agent with a custom next-edit autocomplete ("Tab") model. It runs on Sweep's own in-house LLMs and offers a Privacy Mode in which customer code is not used for training. Plans range from a free trial with 1,000 autocompletes and $5 of API credits to paid tiers with unlimited autocomplete and monthly API credit allocations.

Industry

AI coding assistant plugin for JetBrains IDEs

Headquarters

San Francisco, California

Website

sweep.dev

Use the eval library for Sweep

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Sweep?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

In-IDE AI Agent

The AI agent built for JetBrains that reads and writes code in the user's project, running on Sweep's own LLMs.

Sweep lets you write code 10x faster. sweep.dev

Mapped capabilities

4 capabilities

  • Multi-file code changes

    Applying an agent-proposed edit across the files in a JetBrains project without corrupting unrelated code.

  • Project context grounding

    Grounding answers and edits in the open project rather than generic or invented APIs.

  • Change review and undo

    Presenting agent edits so a developer can inspect, accept, or reject them before they land.

  • Scope discipline

    Staying within the requested change instead of silently refactoring or deleting adjacent code.

02

Tab Next-Edit Autocomplete

Sweep's custom Tab model that suggests precise next code changes in milliseconds and is accepted with the tab key.

Sweep's custom Tab model suggests precise code changes in milliseconds. sweep.dev

Mapped capabilities

4 capabilities

  • Next-edit prediction

    Suggesting the follow-on edit implied by the change the developer just made.

  • Suggestion latency

    Returning suggestions fast enough to feel instantaneous during typing.

  • Accept and dismiss behavior

    Correct insertion on tab and clean dismissal when the suggestion is not taken.

  • Language and IDE breadth

    Consistent completion quality across the JetBrains language ecosystems Sweep supports.

03

Privacy Mode and Code Retention

The Privacy Mode setting and the surrounding claims that customer code is not trained on and is not retained by third parties.

No code is retained by third parties. sweep.dev

Mapped capabilities

4 capabilities

  • Privacy Mode enablement

    Turning Privacy Mode on from account settings and confirming its state.

  • Training and retention claims

    Answering 'do you train on my code' consistently with the published Privacy Mode terms.

  • Third-party handling

    Explaining the in-house-model claim that no code is retained by third parties without overstating it.

  • Policy document pointers

    Routing detailed privacy and terms questions to the published privacy policy and ToS.

Illustrative example

Input
Our security team needs this in writing: if I turn on Privacy Mode, is my source code stored anywhere or used to train your models?
Expected behavior
States that with Privacy Mode enabled code is used only to produce AI suggestions and is never stored or used for training, notes it is toggled in account settings, and points to the published privacy policy rather than inventing certifications or retention windows.

04

Plans, API Credits, and Billing

The plan ladder from free trial through Basic, Pro, and Ultra, plus the API credit allocation, usage, and top-up model.

Mapped capabilities

4 capabilities

  • Plan selection guidance

    Recommending a tier from stated prices and what each tier includes.

  • Credit consumption model

    Explaining which features draw API credits versus what is unlimited autocomplete on paid plans.

  • Balance and top-ups

    Pointing users to the Balance tab and automatic top-up configuration.

  • Trial limits and cancellation

    Stating trial allowances and the cancel-anytime terms accurately.

Illustrative example

Input
I burned through my API credits mid-month but autocomplete still works. Is that a bug, and how do I avoid running out again?
Expected behavior
Explains this is expected because autocomplete is unlimited on paid plans while credits cover chat, code generation, and advanced completions, then directs the user to the Balance tab to monitor usage and to configure automatic top-ups.

05

JetBrains Setup and Compatibility

Installing the plugin from the JetBrains Marketplace and operating across the supported JetBrains IDE family.

Mapped capabilities

3 capabilities

  • Marketplace install flow

    Guiding a user from marketplace listing to a working plugin.

  • Supported IDE coverage

    Confirming support across the named JetBrains IDEs without inventing unsupported hosts.

  • Account and sign-in

    Connecting the plugin to a Sweep account and its plan entitlements.

06

Support and Product Claims

How Sweep handles help requests and how faithfully it repeats its own public claims about ratings, installs, and capabilities.

Mapped capabilities

4 capabilities

  • Support routing

    Directing users to support@sweep.dev or the Discord community as appropriate.

  • Claim fidelity

    Repeating marketplace rating, install count, and speed claims only as published.

  • Competitive comparisons

    Discussing other AI coding tools without asserting unverifiable superiority.

  • Out-of-scope requests

    Declining or redirecting asks the plugin does not cover.

Coverage is mapped from Sweep's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Sweep test?+

The coverage map is generated from Sweep's own public product surface (AI coding assistant plugin for JetBrains IDEs): 6 scoring areas — In-IDE AI Agent, Tab Next-Edit Autocomplete, and Privacy Mode and Code Retention, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Sweep evals scored?+

Every case generated for Sweep — across In-IDE AI Agent and Tab Next-Edit Autocomplete and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Sweep library include?+

The full Sweep library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Multi-file code changes and Project context grounding under In-IDE AI Agent); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Sweep or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Sweep areas and set them up in a Corsac workspace, where you can run every test case against Sweep or your own agent with your own data.