All evals
TC

Eval directory

Evals for Third Chair

Eval coverage for Third Chair, mapped from its public product surface.

About Third Chair

Third Chair is a platform for music rights holders — labels, distributors, publishers, artists, and investment funds — that finds uses of their compositions and recordings across social media and the web. It groups its offering into three solutions: Monitor (detect every use), Enforce (convert unauthorized ads into revenue), and License (close more sync deals). Case studies cite increases in licensing revenue and volumes of unlicensed uses identified for customers such as Symphonic Distribution, Soundstripe, and Duetti.

Industry

music IP monitoring, enforcement, and sync licensing platform

Use the eval library for Third Chair

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Third Chair?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Detection & Attribution (Monitor)

The core scanning capability: finding uses of a client's compositions and recordings across social platforms and the open web, and correctly tying each detected use back to the right work and right rights holder. Grounded in the site's claims of realtime scanning of millions of web pages and per-customer counts of uses identified.

scans millions of web pages in realtime to reveal every use and every opportunity usethirdchair.com

Mapped capabilities

4 capabilities

  • Recall across platforms

    Surfacing uses on the platforms the product claims to monitor, including short-form video and paid placements, without silently dropping a platform from a report.

  • Work-level attribution

    Distinguishing composition from recording, and mapping a detected audio use to the correct title, version, or cover rather than a same-name or similar-sounding work.

  • False-positive discipline

    Declining to flag uses that only superficially resemble the catalog (soundalikes, interpolations without the recording, unrelated works sharing a title).

  • Deduplication of a single use

    Collapsing reposts, crossposts, and multiple crawls of the same asset into one use rather than inflating counts.

02

Use Classification & Licensing Status

The two classification boundaries the FAQ calls out directly: UGC versus advertisement, and licensed versus unlicensed. Every enforcement or licensing action depends on these labels being right, and both have genuine gray zones (creator content that is also a paid partnership; uses covered by a platform blanket deal).

Identify every use of your content, across every platform. usethirdchair.com

Mapped capabilities

4 capabilities

  • UGC vs. paid advertising

    Correctly separating organic fan content from commercial advertising, including branded content and paid partnerships that look organic.

  • Licensed vs. unlicensed determination

    Reasoning about platform blanket licenses, existing direct deals, and territory before calling a use unauthorized.

  • Calibrated uncertainty

    Marking genuinely ambiguous uses as needing human review instead of forcing a confident label.

  • Evidence attached to a label

    Providing the specific observable signals behind a classification so a legal reviewer can check the reasoning.

Illustrative example

Input
A creator video uses our track. Caption reads 'obsessed with this jacket @brand — code MINE20'. No paid-partnership tag. Is this UGC or an ad?
Expected behavior
Should flag this as likely commercial rather than organic, citing the discount code and brand handle as the signals, while noting the missing disclosure tag means the classification is not certain and the use warrants human review before enforcement.

03

Enforcement & Licensing Workflow

Converting classified findings into action: routing unauthorized commercial uses toward monetization and routing demand signals toward sync deals. Reflects the Enforce and License solutions and the reported advertiser response rates.

Third Chair helps licensing teams identify demand, connect with advertisers, and close more sync deals. usethirdchair.com

Mapped capabilities

4 capabilities

  • Triage and prioritization

    Ordering findings by commercial significance rather than raw volume, so teams work the highest-value uses first.

  • Advertiser outreach drafting

    Producing contact material that is accurate about what was found and what is being proposed.

  • Opportunity identification

    Recognizing when a detected use signals sync demand worth pursuing rather than an enforcement matter.

  • State consistency across a matter

    Keeping a use's status coherent as it moves between detected, contested, licensed, and resolved.

05

Reporting & Claim Integrity

The customer-facing numbers the site foregrounds — uses identified, revenue multiples, response rates — and the report surfaces clients act on. This area covers whether the system's own outputs are stated accurately, scoped to what was measured, and free of unsupported extrapolation.

7500+ unlicensed uses identified usethirdchair.com

Mapped capabilities

4 capabilities

  • Metric provenance

    Tying any figure presented to a client back to what was actually counted and over what window.

  • No unsupported extrapolation

    Refusing to project recovered revenue or infringement volume beyond what the detected data supports.

  • Case study accuracy

    Repeating customer results only as published, without transferring one client's outcome onto another's expectations.

  • Coverage caveats

    Stating what a scan did not cover so a client does not read absence of findings as absence of uses.

06

Client Data & Catalog Intake

The onboarding surface implied by the FAQ's "what data do I need to provide" and the free commercial-use audit. Covers ingesting catalog metadata, handling incomplete or messy rights data, and the confidentiality expectations that attach to a rights holder's catalog and license records.

Mapped capabilities

4 capabilities

  • Incomplete catalog handling

    Working from partial metadata and saying what is missing rather than inferring ownership.

  • Ownership and split ambiguity

    Handling co-owned or partially controlled works without overstating what the client can enforce.

  • Client data confidentiality

    Keeping one client's catalog, license terms, and findings out of any other client's context.

  • Audit scope setting

    Being explicit about what the free audit will and will not examine before it runs.

Coverage is mapped from Third Chair's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Third Chair test?+

The coverage map is generated from Third Chair's own public product surface (music IP monitoring, enforcement, and sync licensing platform): 6 scoring areas — Detection & Attribution (Monitor), Use Classification & Licensing Status, and Enforcement & Licensing Workflow, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Third Chair evals scored?+

Every case generated for Third Chair — across Detection & Attribution (Monitor) and Use Classification & Licensing Status and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Third Chair library include?+

The full Third Chair library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Recall across platforms and Work-level attribution under Detection & Attribution (Monitor)); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Third Chair or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Third Chair areas and set them up in a Corsac workspace, where you can run every test case against Third Chair or your own agent with your own data.