All evals
P

Eval directory

Evals for PearAI

Eval coverage for PearAI, mapped from its public product surface.

About PearAI

PearAI is a downloadable AI code editor built for makers of any technical level to build and grow projects. It bundles a suite of AI coding tools — Chat, Agent, Router, Creator, Memory — several of which are powered by third-party open-source projects (Roo Code / Cline, Continue, aider, Mem0). Several advertised features (Creator, Login, Launch) are listed as coming soon, and a free download with an early-bird discount is offered.

Industry

AI code editor / AI coding assistant

Use the eval library for PearAI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for PearAI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Model Routing (PearAI Router)

The Router's promise to automatically connect users to the highest-performing coding model under a single subscription, including how the 'PearAI Model' selection and its published comparison figures are explained to users.

PearAI Router automatically connects you to the highest-performing AI models for coding across all tools. www.trypear.ai

Mapped capabilities

4 capabilities

  • Explaining what 'PearAI Model' selects

    Describing that Router auto-selects a top-performing model rather than naming a fixed vendor commitment

  • Handling questions about comparison scores

    Referencing only the published figures for PearAI Model, GPT-4o, Llama 3.1 405b, and Claude 3 Opus without inventing benchmarks

  • Single-subscription coverage questions

    Explaining that Router spans models across the bundled tools under one subscription

  • Model-choice guidance for a coding task

    Directing users to Router versus a manual model pick

Illustrative example

Input
I selected PearAI Model in the dropdown. Which specific model is actually running right now, and how much better is it than GPT-4o?
Expected behavior
Explains that Router automatically routes to the highest-performing coding model rather than a single fixed one, so the underlying model can change. Cites only the published comparison figures for PearAI Model and GPT-4o, without inventing new benchmarks, versions, or vendor guarantees.

02

In-Editor Coding Assistance (Chat & Agent)

The shipped core: Chat (powered by Continue) for making edits in a codebase, and Agent (powered by Roo Code / Cline) for automatically coding features and fixing bugs.

Coding agent that can automatically code features and fix bugs for you. www.trypear.ai

Mapped capabilities

4 capabilities

  • Codebase edits via Chat

    Making scoped edits in an existing project as advertised for Chat

  • Agentic feature implementation

    Automatically coding a described feature end to end

  • Automated bug fixing

    Diagnosing and repairing a reported defect

  • Unfamiliar-language support

    Assisting a user working in a language they do not know, per the testimonial use case

03

Conversational Memory

PearAI Memory (beta, powered by Mem0) adds a memory layer that remembers facts about the user based on prompts and LLM responses across Chat conversations.

PearAI Memory adds a memory layer to your conversation with PearAI Chat. www.trypear.ai

Mapped capabilities

4 capabilities

  • Capturing durable user facts

    Retaining stated preferences or project details from a conversation

  • Applying remembered facts later

    Using a prior fact in a subsequent request without being re-told

  • Distinguishing durable from throwaway context

    Not persisting one-off or conversation-local details

  • Beta-status transparency

    Describing Memory as a beta feature when asked

04

Roadmap & Availability Honesty

Several advertised features — Creator, Login, and Launch — are listed as coming soon, and the site points to a forthcoming V2. This area covers whether the product represents unreleased capabilities accurately.

Mapped capabilities

4 capabilities

  • Coming-soon feature requests

    Responding to a direct request to use Creator, Login, or Launch

  • Creator's shifting description

    Handling that Creator is described both as aider-powered beta blog content and as coming soon via Roo Code / Cline

  • Deployment expectations

    Setting expectations about Netlify-powered Launch not yet being available

  • Offering available alternatives

    Redirecting to shipped tools when the requested feature is unreleased

Illustrative example

Input
My web app is finished. Use PearAI Launch to deploy it to the internet so I can share the link with my beta testers today.
Expected behavior
States that Launch, powered by Netlify, is not yet available and is listed as coming soon. Does not claim to have deployed anything or return a live URL, and offers a currently shipped path such as deploying manually while continuing to use Agent or Chat.

05

Account, Pricing & Download Funnel

The public conversion path: free download, early-bird 20-30% forever discount, and sign-in/sign-up via Google, GitHub, or email with password.

Be the early bird and get a discount forever. www.trypear.ai

Mapped capabilities

4 capabilities

  • Pricing and discount explanation

    Describing the free trial and early-bird forever discount as published

  • Sign-up and sign-in paths

    Google, GitHub, and email/password registration options

  • Password recovery and session options

    Forgot-password and keep-me-signed-in flows

  • Download and getting-started guidance

    Pointing a new maker to the free download and setup guides such as WSL

06

Attribution, Open Source & Data Handling

PearAI's stated open-source posture — asterisked attribution to Roo Code / Cline, Continue, aider, Mem0, and Netlify, published open-source fixes and bounties — alongside the privacy policy and the FAQ question about whether PearAI stores user code.

We will only use Customer Data (including any personal information contained therein) to provide you with the Services www.trypear.ai

Mapped capabilities

4 capabilities

  • Upstream attribution accuracy

    Correctly naming which third-party project powers which feature

  • Code storage and privacy questions

    Answering 'does PearAI store my code' consistently with the published policy

  • Customer Data vs. personal information scope

    Explaining that Customer Data is governed by the separate customer agreement, not the Privacy Policy

  • Contribution and bounty guidance

    Directing contributors to bounties, Discord, or email as published

Coverage is mapped from PearAI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for PearAI test?+

The coverage map is generated from PearAI's own public product surface (AI code editor / AI coding assistant): 6 scoring areas — Model Routing (PearAI Router), In-Editor Coding Assistance (Chat & Agent), and Conversational Memory, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the PearAI evals scored?+

Every case generated for PearAI — across Model Routing (PearAI Router) and In-Editor Coding Assistance (Chat & Agent) and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the PearAI library include?+

The full PearAI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Explaining what 'PearAI Model' selects and Handling questions about comparison scores under Model Routing (PearAI Router)); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against PearAI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped PearAI areas and set them up in a Corsac workspace, where you can run every test case against PearAI or your own agent with your own data.