All evals
Replit

Eval directory · Code Assistant

Evals for Replit

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Replit AI products.

About Replit

Replit is a browser-based collaborative coding platform; Replit Agent is its autonomous coding agent that turns a prompt into an app plan and builds, iterates, and deploys the full application inside a Repl — wiring Replit Auth, Replit DB, Object Storage, and Autoscale / Reserved VM / Static / Scheduled Deployments, all under a checkpoint-based cost meter.

Employees

~150

Industry

Online IDE & Autonomous Coding Agent

Headquarters

San Francisco, CA

Website

replit.com

Use the eval library for Replit

All 76 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Replit?

9 areas · 76 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ghostwriter Completion Smoke V1

Smoke test for AI coding / IDE assistant completion correctness.

Mapped capabilities

3 scenarios

  • Completion Correctness
  • Style and Maintainability
  • Execution Safety

Example criterion: Replit generates correct, maintainable code completions that satisfy task intent without unsafe patterns.

02

Agent Planning And Build Flow

Evaluates Replit's Agent Planning & Build Flow across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.

Mapped capabilities

9 scenarios

  • prompt-to-plan emission
  • app-spec faithful to prompt
  • 'make it' continuation preserves plan

03

Auth And Replit Auth

Evaluates Replit's Auth & Replit Auth across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.

Mapped capabilities

9 scenarios

  • Replit Auth header trust
  • Replit Auth sign-in button wiring
  • third-party OAuth via Connectors

04

Collaboration And Multiplayer

Evaluates Replit's Collaboration & Multiplayer across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.

Mapped capabilities

9 scenarios

  • concurrent human + Agent edit
  • ownership / share role respect
  • invite link scope

05

Deployments

Evaluates Replit's Deployments across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.

Mapped capabilities

9 scenarios

  • Autoscale vs Reserved VM choice
  • secrets propagation to deployment
  • custom domain TXT/CNAME setup

06

Repl Workspace And Files

Evaluates Replit's Repl Workspace & Files across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.

Mapped capabilities

9 scenarios

  • multi-file edit consistency
  • .replit run command
  • replit.nix package declaration

07

Replit Db And Storage

Evaluates Replit's Replit DB & Storage across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.

Mapped capabilities

9 scenarios

  • Replit DB key namespacing
  • JSON value serialization
  • Object Storage bucket binding

08

Safety Cost And Governance

Evaluates Replit's Safety, Cost & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.

Mapped capabilities

10 scenarios

  • checkpoint cost runaway prevention
  • prompt injection from fetched URL
  • prompt injection from scraped data

09

Tool Use And Midbuild Function Calls

Evaluates Replit's Tool Use & Mid-build Function Calls across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Autonomous Coding Agent eval coverage.

Mapped capabilities

9 scenarios

  • package_install tool grounding
  • run_command output read-back
  • browser_preview iframe URL

Frequently asked questions

What do the Corsac evals for Replit test?+

Each eval pack tests Replit's public product surface — including Ghostwriter Completion Smoke V1, Agent Planning And Build Flow, Auth And Replit Auth — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Replit evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Replit library include?+

The Replit eval library includes 76 graded test cases across 9 eval packs. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Replit or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against Replit or your own agent with your own data.