All evals
LangSmith

Eval directory · AI Platform

Evals for LangSmith

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for LangSmith AI products.

About LangSmith

LangSmith is LangChain's LLM observability and evaluation platform: tracing, datasets, evaluators (LLM-as-judge, code, and human), experiments, prompt management, and online monitoring used by AI teams to measure and improve LLM apps in production.

Employees

~200

Industry

LLM Observability

Headquarters

San Francisco, CA

Use the eval library for LangSmith

All 89 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for LangSmith?

10 areas · 89 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Annotation Queues

Evaluates LangSmith's Annotation Queues across 7 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM observability and evaluation eval coverage.

Mapped capabilities

7 scenarios

  • Queue creation
  • SDK queue management
  • create_feedback

Public sample case

Input
Moderators need queue filtered to feedback.score<0.5 safety runs.
Expected behavior
Create annotation queue in UI or SDK with project scope and filter; route flagged runs; document queue purpose and reviewer RBAC.
Check
Pass / fail check

02

Auth Workspaces Rbac Governance

Evaluates LangSmith's Auth, Workspaces, RBAC & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Observability & Evaluation Platform eval coverage.

Mapped capabilities

10 scenarios

  • workspace-scoped API keys
  • workspace isolation
  • role-based access on projects

Public sample case

Input
Operator generates a LANGSMITH_API_KEY from Settings → API Keys for a CI pipeline. CI must only see one workspace's data.
Expected behavior
Create the key inside the intended workspace; the key is scoped to that workspace by construction and cannot read another workspace's projects or datasets. Store as a CI secret. Rotate on a schedule and on personnel change. Audit via Settings → API Keys → last-used.
Check
Pass / fail check

03

Datasets And Examples

Evaluates LangSmith's Datasets & Examples across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Observability & Evaluation Platform eval coverage.

Mapped capabilities

9 scenarios

  • create_examples bulk ingest
  • split assignment train/test/validation
  • example versioning

Public sample case

Input
Operator uploads 5,000 examples to a new dataset via client.create_examples in a single call.
Expected behavior
Batch upload via create_examples(dataset_id=..., inputs=[...], outputs=[...]) with parallel iterables of equal length. If the SDK / REST imposes a per-call size cap, page in chunks (e.g., 1000) and persist progress so a mid-upload failure can resume by example index. Verify final dataset row count …
Check
Pass / fail check

04

Evaluators

Evaluates LangSmith's Evaluators across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Observability & Evaluation Platform eval coverage.

Mapped capabilities

9 scenarios

  • LLM-as-judge structured output
  • code/heuristic evaluator
  • human/manual feedback

05

Experiments And Comparisons

Evaluates LangSmith's Experiments & Comparisons across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Observability & Evaluation Platform eval coverage.

Mapped capabilities

9 scenarios

  • evaluate() sync orchestration
  • aevaluate() async
  • baseline-vs-candidate comparison

06

Langgraph Platform And Studio

Evaluates LangSmith's LangGraph Platform & Studio across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Observability & Evaluation Platform eval coverage.

Mapped capabilities

9 scenarios

  • deploy a graph to LangGraph Platform
  • threads (long-running conversations)
  • interrupts for human-in-loop

07

Online Monitoring And Feedback

Evaluates LangSmith's Online Monitoring & Feedback across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Observability & Evaluation Platform eval coverage.

Mapped capabilities

9 scenarios

  • threshold alert on error rate
  • latency P99 alert
  • cost spike alert

08

Prompt Hub And Prompt Management

Evaluates LangSmith's Prompt Hub / Prompt Management across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Observability & Evaluation Platform eval coverage.

Mapped capabilities

9 scenarios

  • push_prompt creates commit
  • pull_prompt with commit hash
  • include_model chain pull

09

Tracing And Runs Api

Evaluates LangSmith's Tracing & Runs API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Observability & Evaluation Platform eval coverage.

Mapped capabilities

9 scenarios

  • @traceable decorator parent/child
  • run_type classification
  • distributed tracing headers

10

Workspaces Rbac Billing

Evaluates LangSmith's Workspaces, RBAC & Billing across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM observability and evaluation eval coverage.

Mapped capabilities

9 scenarios

  • API key scoping
  • API key rotation
  • Project isolation

Frequently asked questions

What do the Corsac evals for LangSmith test?+

Each eval pack tests LangSmith's public product surface — including Annotation Queues, Auth Workspaces Rbac Governance, and Datasets And Examples — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the LangSmith evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 89 LangSmith cases — from Auth Workspaces Rbac Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the LangSmith library.

How many test cases does the LangSmith library include?+

The LangSmith eval library includes 89 graded test cases across 10 eval packs, the largest being Auth Workspaces Rbac Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against LangSmith or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 10 LangSmith packs — Annotation Queues and Auth Workspaces Rbac Governance and the rest — against LangSmith or your own agent with your own data.