All evals
LangChain

Eval directory · AI Platform

Evals for LangChain

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for LangChain AI products.

About LangChain

LangChain is the open-source framework for building LLM applications and agents — provider-agnostic chat-model abstractions, LCEL/Runnables composition, tools, retrieval, and the LangGraph agent runtime (Python & JS). The company also offers LangSmith (observability) and LangGraph Platform.

Employees

~200

Industry

Agent Framework

Headquarters

San Francisco, CA

Use the eval library for LangChain

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for LangChain?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agents Langgraph

Evaluates LangChain's Agents (LangGraph) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Orchestration Framework eval coverage.

Mapped capabilities

9 scenarios

  • create_react_agent setup
  • recursion_limit guard
  • checkpointer for thread state

Public sample case

Input
Integrator wants a tool-calling agent and hand-builds a StateGraph from scratch, reimplementing the model/tool loop and introducing routing bugs.
Expected behavior
Use langgraph.prebuilt.create_react_agent(model, tools) for the standard ReAct tool-calling loop; it wires the model node, ToolNode, and conditional routing back to the model. Drop to a custom StateGraph only when the prebuilt loop is insufficient.
Check
Pass / fail check

02

Chat Models And Messages

Evaluates LangChain's Chat Models & Messages across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Orchestration Framework eval coverage.

Mapped capabilities

9 scenarios

  • init_chat_model provider inference
  • message type roles
  • invoke vs stream vs batch

Public sample case

Input
Integrator wants a provider-agnostic chat model and calls init_chat_model('claude-3-5-sonnet-latest') without model_provider, then later swaps to a model id that init_chat_model cannot map to a provider.
Expected behavior
When the model id is unambiguous, rely on init_chat_model's provider inference; when ambiguous, pass model_provider explicitly (e.g., 'anthropic', 'openai'). On an unmappable id, init_chat_model raises — surface the error and require an explicit model_provider rather than defaulting to a hardcoded …
Check
Pass / fail check

03

Lcel And Runnables

Evaluates LangChain's LCEL & Runnables across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Orchestration Framework eval coverage.

Mapped capabilities

9 scenarios

  • pipe operator sequence
  • RunnableParallel fan-out
  • RunnablePassthrough.assign

Public sample case

Input
Integrator composes prompt -> model -> parser by calling each step manually and passing intermediate values by hand instead of using the | operator.
Expected behavior
Compose with the LCEL pipe: chain = prompt | model | StrOutputParser(). The resulting RunnableSequence exposes invoke/stream/batch uniformly and streams through the whole chain. Manual hand-wiring loses streaming, batching, and config propagation.
Check
Pass / fail check

04

Memory And State Langgraph

Evaluates LangChain's Memory & State (LangGraph) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Orchestration Framework eval coverage.

Mapped capabilities

9 scenarios

  • MemorySaver for dev
  • thread_id scoping per user
  • SqliteSaver / PostgresSaver setup

05

Retrieval And Vector Stores

Evaluates LangChain's Retrieval & Vector Stores across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Orchestration Framework eval coverage.

Mapped capabilities

9 scenarios

  • Document page_content + metadata
  • RecursiveCharacterTextSplitter sizing
  • embeddings interface consistency

06

Streaming Callbacks And Safety

Evaluates LangChain's Streaming, Callbacks & Safety across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Orchestration Framework eval coverage.

Mapped capabilities

10 scenarios

  • astream_events token streaming
  • custom callback handler
  • streaming tool-call assembly

07

Structured Output And Parsers

Evaluates LangChain's Structured Output & Parsers across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Orchestration Framework eval coverage.

Mapped capabilities

9 scenarios

  • with_structured_output Pydantic
  • method function_calling vs json_mode
  • PydanticOutputParser format instructions

08

Tools And Tool Calling

Evaluates LangChain's Tools & Tool Calling across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's LLM Orchestration Framework eval coverage.

Mapped capabilities

9 scenarios

  • @tool decorator schema
  • StructuredTool args_schema
  • tool_call_id to ToolMessage pairing

Frequently asked questions

What do the Corsac evals for LangChain test?+

Each eval pack tests LangChain's public product surface — including Agents Langgraph, Chat Models And Messages, and Lcel And Runnables — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the LangChain evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 LangChain cases — from Streaming Callbacks And Safety (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the LangChain library.

How many test cases does the LangChain library include?+

The LangChain eval library includes 73 graded test cases across 8 eval packs, the largest being Streaming Callbacks And Safety with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against LangChain or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 LangChain packs — Agents Langgraph and Chat Models And Messages and the rest — against LangChain or your own agent with your own data.