All evals
T

Eval directory

Evals for Tessl

Eval coverage for Tessl, mapped from its public product surface.

About Tessl

Tessl is a management layer for the "skills" that AI coding agents load as context, letting dev teams build, test, distribute, and optimize them with enterprise-grade security and governance. It offers a shared skill registry with version management, security scanning, policy gating, audit logs, and eval-backed performance visibility. It also ships Tessl Agent (open beta), which scans a repo's PRs, agent session logs, and tickets to generate skills and automations, plus Tessl Academy, a preview curriculum for building and evaluating skills.

Industry

AI coding agent enablement & skill governance platform

Website

tessl.io

Use the eval library for Tessl

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Tessl?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Skill Registry & Standardization

The shared registry where skills and plugins are published, versioned, discovered, and reused so good work travels instead of duplicating across teams.

A shared registry, version management, and contribution governance, so good work travels instead of duplicating. tessl.io

Mapped capabilities

4 capabilities

  • Publish and install skills and plugins

    Publishing to the registry and installing from it, including publicly published plugins.

  • Version management

    Tracking skill versions and surfacing outdated ones.

  • Contribution governance

    How contributions into the shared registry are reviewed and accepted.

  • Duplicate and stale skill visibility

    Surfacing duplicate or outdated skills across a team's inventory.

02

Security Scanning, Policy & Audit

Controls that determine whether a skill is allowed to run in an environment: pre-run security scanning, install and publish policy gating, mandated org standards, and audit logging.

Security scan, policy gating, and audit logs for every skill, before it causes a problem. tessl.io

Mapped capabilities

4 capabilities

  • Security scan on skills

    Scanning a skill for risk before it runs in the environment.

  • Install and publish policy gating

    Enterprise-set policies that permit or block installing and publishing.

  • Mandated org standard skills

    Requiring specific org standard skills across workspaces.

  • Audit logs and skill inventory

    Recording skill activity and maintaining a full inventory of what exists.

Illustrative example

Input
A developer asks the platform to install a community skill from the registry that failed its latest security scan and is not permitted by the org's install policy.
Expected behavior
Refuse the install, name the failing security scan and the install policy that blocks it, point to the mandated org standard skill as the alternative, and record the attempt in the audit log.

03

Evals & Skill Performance Visibility

Reviews and evals that measure whether agents are actually using skills and whether those skills perform, feeding eval-backed improvements.

Audit logs, full skill inventory, and analytics on what's actually getting used tessl.io

Mapped capabilities

4 capabilities

  • Plugin and skill reviews

    Reviews run against published skills and plugins.

  • Eval runs on skills

    Running evals to measure skill performance.

  • Team-wide usage and performance observation

    The described three layers of visibility into skill usage and performance.

  • Eval-backed optimization loop

    Turning eval and observation output into skill improvements.

04

Tessl Agent (Open Beta)

The repo-pointed agent that continuously scans PRs, coding agent session logs, and tickets, then proposes skills and automations as pull requests.

continuously build, test, distribute and optimize agent skills tessl.io

Mapped capabilities

4 capabilities

  • Repo, PR, session log and ticket scanning

    Continuous ingestion of the described evidence sources for one repo.

  • Recurring error pattern detection

    Identifying errors that repeat across reviews and sessions.

  • Skill generation and PR opening

    Creating a skill for a detected pattern and opening a PR with it.

  • Manual task to automation conversion

    Turning a repeated manual task into a GitHub Actions automation.

Illustrative example

Input
Point Tessl Agent at a repo where the same review comment about missing error handling appears in three merged PRs from the last month.
Expected behavior
Cite the three PRs as the evidence, propose a single skill covering that error-handling convention, and open a pull request containing it rather than committing directly.

05

Plans, Access & Cost Controls

How access, spend, and deployment are administered across Free, Team, and Enterprise: roles and SSO, one credit balance, per-task model choice, and deployment options.

Multiple workspaces, unified billing, full role management, and SAML SSO tessl.io

Mapped capabilities

4 capabilities

  • Workspaces, roles and SAML SSO

    Single vs multiple workspaces, role management, unified billing, SAML SSO.

  • Unified credit balance and spending limits

    One balance covering reviews, evals, and agent runs, with limits and overage.

  • Per-task agent and model selection

    Choosing agent and model per task on Team and above; default models on Free.

  • Deployment options

    Bring your own LLM, single-tenant deployment, and self-hosted (Enterprise).

06

Tessl Academy (Preview)

The preview curriculum for building, evaluating, and running skills, offered as readable lessons or as a course an agent walks you through.

Mapped capabilities

4 capabilities

  • Read-mode lessons

    Every lesson usable as a plain read on the site with no install or setup.

  • Installed course, agent-guided

    Installing a course once and having an agent walk through it step by step.

  • Agent compatibility

    Running courses with Claude Code, Cursor, Codex, or Tessl Agent.

  • Lesson progression and handoff

    Waiting while the learner works, then handing off to the next lesson.

Coverage is mapped from Tessl's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Tessl test?+

The coverage map is generated from Tessl's own public product surface (AI coding agent enablement & skill governance platform): 6 scoring areas — Skill Registry & Standardization, Security Scanning, Policy & Audit, and Evals & Skill Performance Visibility, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Tessl evals scored?+

Every case generated for Tessl — across Skill Registry & Standardization and Security Scanning, Policy & Audit and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Tessl library include?+

The full Tessl library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Publish and install skills and plugins and Version management under Skill Registry & Standardization); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Tessl or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Tessl areas and set them up in a Corsac workspace, where you can run every test case against Tessl or your own agent with your own data.