Tool evaluation · 2026
The AI Agent node gets you from nothing to a working demo in an afternoon. Three things break between that demo and production, and all three are documented by n8n itself. Here is where, with the sources.
The fact nobody quotes
Queue mode is how you scale n8n. It disables the AI Agent node’s default memory. That is written in n8n’s own documentation.
In one paragraph
The foundation
n8n's AI Agent node is a layer over LangChain. You attach three kinds of sub-node: a chat model, a memory, and as many tools as you need. It then runs the familiar loop, request, tool call, result, repeat, until the model stops asking for tools.
That is the forty lines you would otherwise write yourself, with error handling and parallel calls included. For a first agent it saves days, and that is a real argument.
One thing to understand before going further: sub-nodes do not behave like normal nodes. A regular node takes n items in and produces n outputs. A sub-node is called by the root node when it needs it, and expression resolution follows different rules. It is the first surprise when you move from a linear workflow to an agent.
The case for it
Self-hosting is a first-class path, not a footnote. This is why regulated sectors pick it. The agent, the workflow data and the logs stay on infrastructure you control, which makes the conversation with a DPO short. The full constraint set is in our guide on GDPR and AI agents.
The integrations already exist. Several hundred of them. An agent that has to read a Google Sheet, post to Slack and write to Postgres is three nodes, not three API clients with their own auth handling and token rotation.
You can see what happened. Every execution shows which node fired, with its input and output. Debugging an agent that picked the wrong tool is a visual exercise rather than a log-reading one. On a project handed to a non-technical client, that difference is worth a lot: they can look themselves.
Credentials live outside the workflow. n8n separates credentials from the workflow body, which avoids the most common leak pattern, an API key hard-coded into a code node.
Break 1
The default memory node is called Simple Memory. It holds a window of conversation history keyed by a session, with a Context Window Length parameter setting how many previous interactions to keep.
Queue mode is exactly what you turn on to scale. In that mode the main instance receives the triggers and passes the execution ID to Redis, which holds the queue; an available worker picks it up, runs it, and writes the result to the database. It is the only way to add capacity by adding machines.
In other words, the mechanism that lets you grow disables the default memory. This is not a bug, it is an architectural consequence, and it is stated in the docs. But it appears in no n8n agent tutorial, because tutorials run in main mode on a single machine.
| Memory node | Where history lives | Survives restart | Queue-mode safe |
|---|---|---|---|
| Simple Memory | Inside the n8n process | No | No, advised against in the docs |
| Postgres Chat Memory | Postgres database | Yes | Yes |
| Redis Chat Memory | Redis | Yes | Yes |
| Motorhead | External service | Yes | Yes |
| Zep | External service | Yes | Yes |
| Xata | External service | Yes | Yes |
The fix is simple: in production, Postgres Chat Memory or Redis Chat Memory. But note what you have even after fixing it: conversation history, not memory. You cannot store "this client refuses Friday calls" as a durable fact, read it back, correct it or audit it. That needs a fact store behind a tool node, which you build yourself. Same conclusion as in what really limits your AI agent.
Break 2
An agent that thinks for a long time is often the one worth having. It is also the one that meets the limits first.
| Limit | Default | Where it is set |
|---|---|---|
| EXECUTIONS_TIMEOUT | -1, meaning no timeout | Environment variable, self-hosted |
| EXECUTIONS_TIMEOUT_MAX | 3,600 seconds | Ceiling on what a user can set per workflow |
| Cloud concurrency | Depends on plan | Not adjustable, excess queued first in first out |
| Process memory | Depends on the machine | NODE_OPTIONS, or add workers |
Running out of memory is the nastiest failure mode, because it does not always announce itself. The documentation lists the messages to recognise: "Execution stopped at this node (n8n may have run out of memory while executing it)", along with "Problem running workflow", "Connection Lost", "503 Service Temporarily Unavailable", and server-side "Allocation failed - JavaScript heap out of memory".
On Cloud and on the official Docker image, n8n restarts automatically when this happens. On a hand-rolled install it does not: you find out when the client calls. That is an argument for monitoring before go-live, covered in deploying an AI agent to production.
Why does an agent consume so much? Because every turn resends the whole conversation and tool results pile up inside it. A tool that returns two hundred database rows injects those rows into the context, and they stay there for every subsequent turn. Same trap as on the MCP side, detailed in MCP in the enterprise.
Break 3
An n8n workflow is a JSON document. While it describes "when a form arrives, write a row to a spreadsheet", the absence of version control is tolerable. Once it describes an agent deciding on client files, it becomes a governance problem before it is a technical one.
Below that tier, manual JSON export is what you have. Technically you can commit it. In practice nobody reviews a generated-JSON diff, and two people cannot approve a change the way they approve a pull request. For an agent touching legal, medical or HR data, this is the question your first audit will ask.
Credit where due
The standard complaint about n8n was that you could not test an agent. That is no longer true and it deserves saying.
n8n shipped an evaluations feature. The principle is what you would expect: you assemble a test dataset, each case with a sample input and often an expected output, run it through the workflow, and metrics measure answer quality. Their documentation frames it as "the difference between a flaky proof of concept and a solid production workflow", and recommends using it during building and after deploying.
What it covers: the quality of the agent's output. What it does not: change review, regression when an integration changes behaviour, and infrastructure. Real progress, not a substitute for version control.
The real numbers
| Plan | Price | Executions | Git versioning | Who it fits |
|---|---|---|---|---|
| Community, self-hosted | Free, you pay the server | Unlimited | No | Regulated data, or high volume |
| Starter | 20 EUR / month | 2,500 | No | A first agent, one team |
| Pro | 50 EUR / month | 10,000 | No | Production, small team |
| Business | 667 EUR / month | 40,000 | Yes | Under 100 employees, governance required |
| Enterprise | On request | Negotiated | Yes | SSO, network isolation, support contract |
Read the third and fourth rows together. Going from Pro to Business is thirteen times the price for four times the executions. What you are actually buying in that gap is not executions: it is collaboration and version control.
Which is why most teams crossing 10,000 executions a month move to self-hosting rather than pay the tier. It is also the most common request we get about n8n, ahead of any technical question.
The forgotten line
One agent run counts as one execution, however many times it loops internally. That is generous billing on n8n's part, and it deserves acknowledging.
But the cost that follows the loop is elsewhere. A five-turn agent is five calls to the model API, each resending the whole conversation accumulated so far. n8n shows this nowhere: you find it on your model provider's invoice, the following month.
Two levers do most of the work. Caching the system prompt and tool definitions, which are identical every turn. And not letting a tool dump hundreds of rows into the context: paginate, and give the agent a second tool to fetch the detail of a single record. Per-model rates and the full arithmetic are in Claude API pricing for business.
The misunderstanding
n8n is commonly called open source. It is not, in the OSI sense. It ships under a Sustainable Use License, a fair-code model: the source is readable and modifiable, internal use is free, but reselling it as a hosted service is forbidden.
For a company using it for itself, the constraint is theoretical. For an agency wanting to resell hosted n8n to its clients, it is not. And for a public body or a large account whose internal policy requires an OSI-approved licence, it is grounds for rejection in committee. The detailed consequences, and nine genuinely open alternatives, are in the best n8n alternatives by licence.
Our position
| Situation | Our choice | Why |
|---|---|---|
| Agent with under 10 tools, one job | n8n | Built in days, debuggable by the client |
| Regulated data, on-premise required | n8n self-hosted | The data never leaves their infrastructure |
| Scaling out with queue mode | n8n, Postgres memory | Simple Memory is advised against in that mode |
| Agent needs durable, auditable memory | Code, with n8n for the plumbing | A fact store you can read and correct |
| Decisions on client files, review required | Code | n8n versioning starts at 667 EUR a month |
| Internal policy requires an OSI licence | An alternative | The Sustainable Use License is not OSI |
In practice most of our deployments are hybrid. n8n carries the triggers, the integrations, the retries and the visual monitoring. The agent logic and its memory live in code, behind a single HTTP tool node. You keep visual debugging where it helps and version control where it matters, and you do not pay 667 EUR a month to put a JSON file in git.
That is what we build as custom AI agents, and how we work as an n8n agency. For what that looks like in practice, our 12 workflow examples by industry give the real running costs.
FAQ
Related guides
What you are legally allowed to do with each, which is the real reason teams leave.
Concrete builds with their real running costs.
The standard way to give an agent tools, and how to keep it fast.
Monitoring, rollback and the human checkpoint.
Links verified at publication. Regulatory texts change — always defer to the official source.
A question, a project, an idea? We respond within 24h. Free audit, no commitment.