All evals
Linear

Eval directory · Code Assistant

Evals for Linear

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Linear AI products.

About Linear

Linear is a project and issue management tool built for high-velocity software teams. It pairs a fast, keyboard-driven UI with a GraphQL API, cycles, projects, and roadmaps, plus deep GitHub and Slack integrations and AI-assisted triage.

Employees

~100

Industry

Developer Productivity

Headquarters

San Francisco, CA

Website

linear.app

Use the eval library for Linear

All 78 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Linear?

8 areas · 78 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ai Triage

Evaluates Linear's AI Triage & Assistance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Project & Issue Management eval coverage.

Mapped capabilities

10 scenarios

  • Suggested assignee grounding
  • Label suggestion evidence
  • Summary citation to issue body

Public sample case

Input
Issue text mentions Payments API errors. Team Payments on-call rotation exists. AI suggestion must cite team and past similar issues, not random user.
Expected behavior
Propose on-call from Payments team with citation to issue content; allow operator override; tag model confidence [REQUIRES-VERIFICATION].
Check
Pass / fail check

02

Api Graphql

Evaluates Linear's API & GraphQL Contract across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Project & Issue Management eval coverage.

Mapped capabilities

10 scenarios

  • Cursor pagination (first/after)
  • Complexity-aware queries
  • RATELIMITED backoff

Public sample case

Input
issues connection supports first/after with pageInfo.endCursor at api.linear.app/graphql. Operator needs complete open-issue set for audit export.
Expected behavior
Loop GraphQL queries with first (<=50 default cost) and after=cursor until pageInfo.hasNextPage is false; never client-filter an oversized single fetch.
Check
Pass / fail check

03

Automation Rules

Evaluates Linear's Automation Rules across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Project & Issue Management eval coverage.

Mapped capabilities

9 scenarios

  • Trigger condition safety
  • Side-effect idempotency
  • Label automation loops

Public sample case

Input
Workspace automation triggers on Issue label added. Condition must scope to team Infosec to avoid firing on unrelated teams' issues.
Expected behavior
Automation filter includes team + label security; assignee set to on-call rotation user; test with sample issue before enable.
Check
Pass / fail check

04

Cycles Projects Roadmap

Evaluates Linear's Cycles, Projects & Roadmap across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Project & Issue Management eval coverage.

Mapped capabilities

10 scenarios

  • Cycle capacity rollups
  • Project scope changes
  • Initiative progress accuracy

05

Import Export Data

Evaluates Linear's Import/Export & Data Integrity across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Project & Issue Management eval coverage.

Mapped capabilities

9 scenarios

  • CLI importer fidelity
  • Attachment URL handling
  • Deduplication on import

06

Issue Lifecycle

Evaluates Linear's Issue Lifecycle & State Transitions across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Project & Issue Management eval coverage.

Mapped capabilities

10 scenarios

  • Triage queue routing
  • Workflow stateId integrity
  • Sub-issue parent linkage

07

Permissions Workspace

Evaluates Linear's Permissions & Workspace Access across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Project & Issue Management eval coverage.

Mapped capabilities

10 scenarios

  • Guest team scoping
  • Admin API key governance
  • OAuth actor authorization

08

Webhooks Integrations

Evaluates Linear's Webhooks & Integrations across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Project & Issue Management eval coverage.

Mapped capabilities

10 scenarios

  • Webhook payload previousValues
  • GitHub PR issue linking
  • Slack issue creation

Frequently asked questions

What do the Corsac evals for Linear test?+

Each eval pack tests Linear's public product surface — including Ai Triage, Api Graphql, and Automation Rules — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Linear evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 78 Linear cases — from Ai Triage (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Linear library.

How many test cases does the Linear library include?+

The Linear eval library includes 78 graded test cases across 8 eval packs, the largest being Ai Triage with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Linear or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Linear packs — Ai Triage and Api Graphql and the rest — against Linear or your own agent with your own data.