All evals
K

Eval directory

Evals for Klarity

Eval coverage for Klarity, mapped from its public product surface.

About Klarity

Klarity is an AI transformation platform that builds "Context Graphs" of how work actually gets done inside an enterprise. It runs three phases — Discover (passive screen observation, AI interviews, and ingestion of existing SOPs), Structure (a searchable, continuously updated Process Index), and Improve (an Advisor that flags bottlenecks and automation ROI, plus Signals that surface team practices). It is positioned at finance and operations transformation, with named customers including DoorDash and Tyler Technologies.

Industry

enterprise process intelligence / AI transformation platform

Use the eval library for Klarity

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Klarity?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Discovery capture

Phase 1: turning passive screen observation, AI interviews, recorded walkthroughs, and existing documentation into structured process records without integrations.

Ingests your existing SOPs, process docs, and runbooks instantly. www.klarity.ai

Mapped capabilities

4 capabilities

  • Screen-activity to process steps

    Converting passively observed app activity into ordered, deduplicated workflow steps.

  • Recorded walkthrough to usable SOP

    Producing a publishable SOP from a video or transcript, as described in the DoorDash workflow.

  • AI Interviewer elicitation

    Structured conversations that surface judgment calls, workarounds, and exceptions observation alone misses.

  • Existing document intake

    Ingesting SOPs, process docs, and runbooks and reconciling them against observed reality.

Illustrative example

Input
A Zoom transcript of an analyst walking through invoice approval, including a spoken aside that invoices over $50K also need the controller's sign-off. Produce an SOP.
Expected behavior
The SOP lists the approval steps in the order demonstrated and represents the $50K controller sign-off as a conditional branch rather than a flat step. No step appears that was not shown or stated in the transcript.

02

Process Index structure

Phase 2: organizing discovery output into a searchable, role-scoped index that stays current as workflows change.

The Process Index updates as your reality changes. www.klarity.ai

Mapped capabilities

4 capabilities

  • Process mapping and grouping

    Assembling steps into named end-to-end processes such as quote-to-cash.

  • Role-scoped search and retrieval

    Returning the processes relevant to a requester's function rather than the full corpus.

  • Continuous currency

    Detecting that a workflow has changed and updating the indexed process instead of stale duplication.

  • Duplicate and variant handling

    Distinguishing a genuine process variant from a redundant re-capture at scale.

03

Advisor and Signals

Phase 3: surfacing bottlenecks, automation opportunities ranked by ROI, and team practices worth scaling.

Mapped capabilities

4 capabilities

  • Bottleneck and redundancy identification

    Flagging where a mapped process stalls or repeats work.

  • ROI-prioritized automation opportunities

    Ranking candidates and showing the basis for each projected figure.

  • Signals practice detection

    Identifying an individual's better method and assessing whether it generalizes to others.

  • Executive-level summarization

    Rolling findings into a decision-ready view for a CAO or operating partner.

Illustrative example

Input
Three indexed processes with volumes and cycle times supplied, and a fully loaded hourly rate. Rank them as automation opportunities by projected annual ROI.
Expected behavior
All three are ranked with a projected annual ROI computed from the supplied volumes, cycle times, and rate, with the assumed automation rate stated. No dollar figure or benefit is introduced that the supplied fields cannot produce.

04

Context Graph fidelity

The reasoning layer Klarity claims over a plain process map: decision traces, exceptions, and tribal knowledge, faithfully attributed to their source.

Mapped capabilities

4 capabilities

  • Decision and judgment-call capture

    Recording why a path was taken, not just that it was.

  • Exception and edge-path representation

    Preserving the non-happy-path branches that make operations actually run.

  • Provenance and traceability

    Tracing each captured element back to the observation, interview, or document it came from.

  • Uncertainty and gap marking

    Declaring what remains unknown rather than smoothing over an incomplete picture.

05

Observation governance

The privacy, consent, and access surface implied by continuous passive screen capture inside an enterprise, and by Klarity's stated privacy commitments.

Mapped capabilities

4 capabilities

  • Sensitive content handling in capture

    Behavior when observed screens contain personal, financial, or credential data.

  • Consent and transparency to observed employees

    What the observed person is told and can see about their own capture.

  • Access scoping across roles and units

    Keeping a team's processes visible only to those entitled to them.

  • Privacy-rights request handling

    Responding to state privacy requests consistent with the published CCPA position.

06

Transformation workflows

The multi-week engagements the case studies describe: post-M&A integration, onboarding acceleration, and continuous rather than episodic improvement.

A complete current-state picture in days. www.klarity.ai

Mapped capabilities

4 capabilities

  • Multi-business-unit current-state assembly

    Consolidating many units into one coherent picture, as in the Tyler Technologies integration.

  • Onboarding and training enablement

    Turning indexed processes into material a new hire can actually follow.

  • Program progress and coverage reporting

    Reporting how much of an operation has been captured and what is still dark.

  • Change tracking over time

    Showing what shifted between two points in a continuous program.

Coverage is mapped from Klarity's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Klarity test?+

The coverage map is generated from Klarity's own public product surface (enterprise process intelligence / AI transformation platform): 6 scoring areas — Discovery capture, Process Index structure, and Advisor and Signals, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Klarity evals scored?+

Every case generated for Klarity — across Discovery capture and Process Index structure and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Klarity library include?+

The full Klarity library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Screen-activity to process steps and Recorded walkthrough to usable SOP under Discovery capture); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Klarity or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Klarity areas and set them up in a Corsac workspace, where you can run every test case against Klarity or your own agent with your own data.