All evals
n8n

Eval directory · AI Platform

Evals for n8n

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for n8n AI products.

About n8n

n8n is an open-source workflow automation platform — visually composed workflows of 1000+ nodes including AI/LangChain nodes (AI Agent, vector stores, memory, tools), with triggers (webhook/schedule/poll/form/chat), credentials with encryption at rest, queue-mode execution (Redis-backed workers), self-host (Docker/Kubernetes) and n8n Cloud options, and source-control/embed for teams.

Employees

~100

Industry

Workflow Automation

Headquarters

Berlin, Germany

Website

n8n.io

Use the eval library for n8n

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for n8n?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ai Langchain Nodes

Evaluates n8n's AI / LangChain Nodes across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Workflow Automation eval coverage.

Mapped capabilities

9 scenarios

  • AI Agent tool wiring
  • JSON output parser
  • vector store retriever k

Public sample case

Input
Operator drops an AI Agent node, an OpenAI Chat Model node, and three HTTP Request Tool nodes onto the canvas but only connects the Chat Model. The agent responds without ever calling the tools.
Expected behavior
Each tool node must be connected to the AI Agent's `ai_tool` input (the dotted-line port labelled 'Tool'). Verify the agent's exposed tools list reflects all three. Tool nodes provide their `name` + `description` to the LLM — confirm both are set so the model can route.
Check
Pass / fail check

02

Credentials And Auth

Evaluates n8n's Credentials & Auth across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Workflow Automation eval coverage.

Mapped capabilities

9 scenarios

  • N8N_ENCRYPTION_KEY rotation
  • OAuth2 redirect URI mismatch
  • OAuth2 scope minimisation

Public sample case

Input
Ops team rotates N8N_ENCRYPTION_KEY in the docker-compose env without re-encrypting existing credentials. All workflows fail with 'cannot decrypt credential'.
Expected behavior
Rotation is a two-step playbook: run `n8n executeBatch --re-encrypt` (or the documented re-encryption procedure) with both old and new keys configured BEFORE removing the old key. Test on a staging instance. Back up the DB first.
Check
Pass / fail check

03

Execution Engine Queue Mode And Scaling

Evaluates n8n's Execution Engine, Queue Mode & Scaling across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Workflow Automation eval coverage.

Mapped capabilities

9 scenarios

  • regular vs queue mode
  • worker crash mid-execution
  • EXECUTIONS_TIMEOUT

Public sample case

Input
Operator runs n8n in a single Docker container under load. Workflow latency spikes because the main process is both serving the UI and running long workflows.
Expected behavior
Switch to EXECUTIONS_MODE=queue with a Redis URL and one or more `n8n worker` processes. Main process handles UI + trigger registration; workers consume jobs. Pin worker concurrency via N8N_CONCURRENCY_PRODUCTION_LIMIT per worker. Do not leave both modes mixed across instances.
Check
Pass / fail check

04

Safety Rbac Compliance And Governance

Evaluates n8n's Safety, RBAC, Compliance & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Workflow Automation eval coverage.

Mapped capabilities

10 scenarios

  • global roles owner/admin/member
  • SAML / OIDC SSO enforcement
  • audit log retention

05

Self Host Cloud And Embed

Evaluates n8n's Self-host, Cloud & Embed across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Workflow Automation eval coverage.

Mapped capabilities

9 scenarios

  • Docker volume persistence
  • reverse proxy + WEBHOOK_URL
  • K8s pod scaling vs single PV

06

Triggers

Evaluates n8n's Triggers across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Workflow Automation eval coverage.

Mapped capabilities

9 scenarios

  • test URL vs production URL
  • webhook auth=Header / Basic / JWT
  • responseMode=onReceived vs lastNode

07

Workflows And Nodes

Evaluates n8n's Workflows & Nodes across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Workflow Automation eval coverage.

Mapped capabilities

9 scenarios

  • items[] is array, not single object
  • $json vs $node[name].json scoping
  • continueOnFail vs error output

08

Workflows As Code And Api

Evaluates n8n's Workflows-as-Code & API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Workflow Automation eval coverage.

Mapped capabilities

9 scenarios

  • REST API key auth header
  • workflow export/import via CLI
  • source control + environments

Frequently asked questions

What do the Corsac evals for n8n test?+

Each eval pack tests n8n's public product surface — including Ai Langchain Nodes, Credentials And Auth, and Execution Engine Queue Mode And Scaling — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the n8n evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 n8n cases — from Safety Rbac Compliance And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the n8n library.

How many test cases does the n8n library include?+

The n8n eval library includes 73 graded test cases across 8 eval packs, the largest being Safety Rbac Compliance And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against n8n or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 n8n packs — Ai Langchain Nodes and Credentials And Auth and the rest — against n8n or your own agent with your own data.