
Eval directory
Evals for LangSmith
10 evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for LangSmith AI products.
About LangSmith
LangSmith is LangChain's LLM observability and evaluation platform: tracing, datasets, evaluators (LLM-as-judge, code, and human), experiments, prompt management, and online monitoring used by AI teams to measure and improve LLM apps in production.
Employees
~200
Industry
LLM Observability
Headquarters
San Francisco, CA
Website
www.langchain.com/langsmithHow complete this published benchmark library is across datasets, metrics, rubrics, use-case maps, and pack context. This is library coverage, not an agent performance score.
Test datasets
10/10 packs
Scoring metrics
0/10 packs
Judge rubrics
10/10 packs
Use-case maps
0/10 packs
Pack context
10/10 packs
Available eval packs for LangSmith
10 packs ready to run.
Annotation Queues
Evaluates LangSmith's Annotation Queues across 7 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM observability and evaluation eval coverage.
Auth Workspaces Rbac Governance
Evaluates LangSmith's Auth, Workspaces, RBAC & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Observability & Evaluation Platform eval coverage.
Datasets And Examples
Evaluates LangSmith's Datasets & Examples across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Observability & Evaluation Platform eval coverage.
Evaluators
Evaluates LangSmith's Evaluators across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Observability & Evaluation Platform eval coverage.
Experiments And Comparisons
Evaluates LangSmith's Experiments & Comparisons across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Observability & Evaluation Platform eval coverage.
Langgraph Platform And Studio
Evaluates LangSmith's LangGraph Platform & Studio across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Observability & Evaluation Platform eval coverage.
Online Monitoring And Feedback
Evaluates LangSmith's Online Monitoring & Feedback across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Observability & Evaluation Platform eval coverage.
Prompt Hub And Prompt Management
Evaluates LangSmith's Prompt Hub / Prompt Management across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Observability & Evaluation Platform eval coverage.
Tracing And Runs Api
Evaluates LangSmith's Tracing & Runs API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Observability & Evaluation Platform eval coverage.
Workspaces Rbac Billing
Evaluates LangSmith's Workspaces, RBAC & Billing across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM observability and evaluation eval coverage.
Why eval LangSmith AI
LangSmith's AI features ship behind brand promises about accuracy, safety, and reliability. Buyers and integrators need to know those promises hold up under adversarial prompts, edge-case workflows, and the long tail of real customer inputs — not just the demo path.
The Corsac eval library for LangSmith measures four dimensions teams care about most when deploying ai platform agents:
- Adversarial robustness — does the agent resist prompt injection, jailbreaks, and social-engineering attempts?
- Workflow quality— does it complete the task buyers were shown in the demo, on inputs that don't look like the demo?
- Safety gates — does it escalate or refuse when it should, and only then?
- Operator quality — does it preserve analyst trust by surfacing the right context at the right time?
Every eval pack above is hand-authored against LangSmith's public product surface and runnable in Corsac with your own data.