Architecture · 2026
Almost every team that says “our agent is self-hosted” is still sending every prompt and every document to someone else’s API. That is not a criticism, it is usually the right call. But you should know which one you bought.
The sentence that reframes the whole project
Self-hosting the agent and self-hosting the model are two different decisions. One is free. The other starts at 574.87 EUR a month before you process a single request.
In one paragraph
The model
This is the distinction that resolves most arguments about self-hosting, and it is rarely drawn explicitly.
| Layer | What it does | Can you host it? | Cost of hosting it |
|---|---|---|---|
| Orchestration | The loop, the triggers, the integrations | Yes, easily | A small VM, tens of euros |
| Model | The reasoning | Yes, with a GPU | From 574.87 EUR a month |
| Data | Your documents, your database | Yes, you already do | Already in your budget |
Hosting layers one and three is cheap and almost always worth doing. Hosting layer two is an order of magnitude more expensive and is the one people mean when they say "self-hosted AI".
The expensive misunderstanding
The most common build we are asked to review looks like this: n8n installed on the company's own server, workflows in their own database, logs in their own infrastructure, and an AI Agent node pointing at a hosted model API.
Everything about that is reasonable. The orchestration genuinely is self-hosted, which is a real security and compliance benefit, covered in our evaluation of n8n for AI agents.
Teams discover this at the worst possible moment, which is during a security review or a client questionnaire, after the system is live. The fix at that point is architectural, not a setting.
Be precise
| Architecture | Prompts leave? | Documents leave? | Logs leave? | Monthly floor |
|---|---|---|---|---|
| Hosted platform, hosted model | Yes | Yes | Yes | Platform subscription |
| Self-hosted orchestration, hosted model | Yes | Yes | No | A small VM |
| Self-hosted orchestration, hosted model, ZDR agreement | Yes, not retained | Yes, not retained | No | A small VM |
| Self-hosted orchestration, self-hosted model | No | No | No | 574.87 EUR |
Note the third row. It is the one most teams should be in, and the one almost nobody knows exists.
The lever nobody pulls
Anthropic's commercial data handling page, updated 1 July 2026, states that for API users inputs and outputs are automatically deleted from the backend within 30 days of receipt or generation. It lists the exceptions plainly: services with longer retention under your own control such as the Files API, legal obligations, usage policy enforcement, and cases where you and they have agreed otherwise.
For a legal, HR or healthcare workload this changes the conversation materially. Thirty-day retention is often acceptable with the right paperwork; zero retention removes the objection almost entirely, and it costs nothing but a conversation with a sales team.
It does not remove the transfer question, the processor agreement, or your own obligations. The full set of what to get in writing, provider by provider, is in GDPR and AI agents.
The real number
This is where self-hosting stops being free. A model needs a GPU, and a GPU needs to be running whether or not anyone is using it.
| Instance | GPUs | RAM | Per hour | Per month, continuous |
|---|---|---|---|---|
| L4-1-24G | 1 | 48 GB | 0.79 EUR | ~574.87 EUR |
| L4-2-24G | 2 | 96 GB | 1.58 EUR | ~1,149.75 EUR |
| L4-4-24G | 4 | 192 GB | 3.15 EUR | ~2,299.50 EUR |
| L4-8-24G | 8 | 384 GB | 6.30 EUR | ~4,599 EUR |
Read the first row carefully, because it is the floor and not a typical configuration. 574.87 EUR a month buys one 24 GB card, which runs small or quantised models comfortably and the current frontier reasoning models not at all. Matching the quality your team expects usually means the third or fourth row.
And that is before the parts nobody budgets: someone to keep the inference server running, the model updates, the capacity planning for concurrent requests, and the fact that an idle GPU bills exactly like a busy one. Serving is usually done with vLLM or Ollama, both of which are free and neither of which runs itself.
Compare against the API side honestly. For the volumes a mid-size company generates, token costs typically land well below the GPU floor, and prompt caching lowers them further. The arithmetic is laid out in Claude API pricing for business.
When to pay it
There are cases where the GPU bill is simply the cost of being allowed to operate.
| Situation | Why hosted will not do |
|---|---|
| Health data under HDS certification | The hosting itself is certified and audited; an API call leaves that perimeter |
| Legal professional secrecy | The duty is personal to the lawyer and does not transfer to a processor by contract alone |
| Defence, sovereignty, classified | The requirement is usually network isolation, not a contract clause |
| Client contract forbidding third-party processing | Written into the deal, not negotiable by you |
| Very high, steady volume | At enough tokens per month the GPU is genuinely cheaper |
The first two are the ones we see in practice. The health case has its own architecture, covered in AI agents and health data under HDS. The legal case is subtler than most vendors admit and we wrote it up in AI in law firms and professional secrecy.
Note what is not on that list: "our data is sensitive", "we are a European company", and "the board is nervous about AI". Those are real concerns and they are usually answered by the third row of the earlier table, not by buying GPUs.
The uncomfortable part
Most requests we receive for a fully self-hosted agent are driven by a feeling rather than a constraint. That feeling is legitimate and the answer to it is usually cheaper than a GPU.
Three questions settle it. Is there a written rule, in a contract, a certification or a regulation, that forbids a processor? Would a zero data retention agreement satisfy whoever is worried? Does your monthly token volume exceed the GPU floor? If all three answers are no, you are about to pay 575 EUR a month and accept a weaker model to solve a problem you do not have.
Saying this costs us work, because the fully self-hosted build is the larger project. We say it because the alternative is a system that is expensive, slower, and abandoned in eighteen months.
Our position
For most clients in regulated sectors, the architecture that survives contact with a security review looks like this:
That is what we build as custom AI agents, and how we work as an n8n agency. The production checklist that goes with it is in deploying an AI agent to production, and the security surface in AI agent security.
FAQ
Related guides
The orchestration layer, evaluated honestly.
What to get in writing from each provider before you send anything.
The one sector where the hosting itself is certified.
What an agent can reach, and what that means when it is wrong.
Links verified at publication. Regulatory texts change — always defer to the official source.
A question, a project, an idea? We respond within 24h. Free audit, no commitment.