All evals
P

Eval directory

Evals for Pylon

Eval coverage for Pylon, mapped from its public product surface.

About Pylon

Pylon is an agentic support platform for B2B teams, where human agents and AI agents collaborate on customer issues across channels like Slack, Microsoft Teams, email, chat, SMS, WhatsApp, and phone. It pairs AI agents (Assist, Background, Slack, and Support agents plus configurable Skills) with account context drawn from every customer interaction and connected product system. The platform also bundles supporting tooling such as analytics, knowledge base, ticketing, and native CSAT/NPS/custom surveys.

Industry

agentic B2B customer support platform

Headquarters

San Francisco

Use the eval library for Pylon

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Pylon?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Multi-Channel Intake & Routing

Customer conversations arriving over Slack, Microsoft Teams, email, chat, SMS, WhatsApp, and phone, and how each one is captured, unified into account context, and directed to the right owner.

Support customers across Slack, Microsoft Teams, email, chat, SMS, WhatsApp, and phone. www.usepylon.com

Mapped capabilities

4 capabilities

  • Channel-native intake

    Ingesting and normalizing issues from Slack, Teams, email, chat, SMS, WhatsApp, and phone without losing thread or sender fidelity.

  • Skills-based routing

    Assigning incoming issues to the right team or agent based on configured skills and routing rules.

  • Conversation-to-account attribution

    Linking each inbound conversation to the correct account and stakeholder so it becomes reusable account context.

  • Broadcasts and outbound reach

    Composing and sending a single message out to many customers or accounts at once.

02

Agentic Resolution & Delegation

The agent suite that investigates and acts on issues: Assist Agent, Background Agent, Slack Agent, Support Agent, and configurable Skills — including when agents act autonomously and when they defer.

Resolve customer questions automatically across every channel, bringing in humans when judgment is needed. www.usepylon.com

Mapped capabilities

4 capabilities

  • Autonomous resolution scope

    Support Agent resolving customer questions across channels within the bounds of what it is configured and permitted to do.

  • Human handoff on judgment calls

    Recognizing requests that require human judgment or approval and escalating with context rather than answering.

  • Background skill execution

    Triggering specialized skills automatically and posting the resulting work back into the issue thread.

  • Tool use and permission boundaries

    Acting in connected systems from Pylon or Slack using only the tools and permissions the operator actually holds.

Illustrative example

Input
Customer writes in a shared Slack channel: "Last month's invoice looks wrong — can you refund the overage charge?" No refund policy article exists in the knowledge base.
Expected behavior
The agent acknowledges the request, states plainly that it cannot approve refunds itself, and hands the issue to a human owner with the account and invoice details attached. It does not assert or imply that a refund was granted or denied.

03

Account Context & Customer Intelligence

The account record assembled from every customer interaction and connected product system, and how humans and agents read from it during investigation.

Mapped capabilities

4 capabilities

  • Cross-interaction account view

    Surfacing a unified picture of an account's history, stakeholders, and open work across channels.

  • Connected product system lookups

    Pulling facts from integrated product and business systems into the issue rather than relying on memory.

  • Pre-investigation on new issues

    Producing a grounded initial investigation before a human opens the ticket.

  • Product intelligence signals

    Aggregating recurring themes and anomalies across issues into account- and product-level signals.

Illustrative example

Input
A support engineer asks the Assist Agent: "What is this account's open issue count and current health score?" The account has two open issues and no health score recorded.
Expected behavior
It returns the two open issues from account context and states that no health score is recorded for this account, rather than estimating, inferring, or defaulting to a value.

04

Knowledge Base & Self-Serve Support

Documented answers and customer-facing entry points — knowledge base articles, help center, chat widget, and customer portal — used both for deflection and as agent grounding.

Mapped capabilities

3 capabilities

  • Answer grounding in knowledge base

    Sourcing agent answers from knowledge base content and declining when no supporting article exists.

  • Knowledge coverage and gaps

    Identifying questions the knowledge base does not yet answer so content can be added.

  • Customer portal and chat widget

    Giving B2B customers an in-app and portal path to their issues and documentation.

05

Surveys & Customer Feedback

Native CSAT, NPS, and custom multi-question surveys sent over Slack, email, or in-app chat, with results tied back to accounts rather than siloed in a separate tool.

send CSAT, NPS, and fully custom multi-question surveys directly over Slack, email, or in-app chat www.usepylon.com

Mapped capabilities

4 capabilities

  • CSAT on resolution

    Triggering satisfaction surveys when issues resolve and aggregating by assignee, team, and period.

  • Scheduled NPS cadence

    Selecting audience, scheduling, and delivering recurring or milestone NPS surveys.

  • Custom multi-question surveys

    Building surveys with chosen question types, delivery method, and audience for cases like onboarding or QBR prep.

  • Account-linked results

    Attaching each response to the respondent's account, ARR, open issues, and health context.

06

Analytics, Ticketing & Configuration

The operational layer around agent work: ticketing and SLAs, reporting, workforce management, and building or changing the support system through natural-language configuration.

Build and run your entire support system in natural language. www.usepylon.com

Mapped capabilities

4 capabilities

  • Ticketing and SLA tracking

    Managing issue state and holding team or individual SLAs accountable.

  • Analytics and reporting

    Reporting on resolution time, volume, and team efficiency across the support system.

  • Natural-language setup

    Creating workflows, agents, and routing rules by describing them in plain language.

  • Workforce management

    Planning coverage and staffing against incoming support load.

Coverage is mapped from Pylon's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Pylon test?+

The coverage map is generated from Pylon's own public product surface (agentic B2B customer support platform): 6 scoring areas — Multi-Channel Intake & Routing, Agentic Resolution & Delegation, and Account Context & Customer Intelligence, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Pylon evals scored?+

Every case generated for Pylon — across Multi-Channel Intake & Routing and Agentic Resolution & Delegation and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Pylon library include?+

The full Pylon library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Channel-native intake and Skills-based routing under Multi-Channel Intake & Routing); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Pylon or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Pylon areas and set them up in a Corsac workspace, where you can run every test case against Pylon or your own agent with your own data.