Architecture · 2026

Self-hosted AI agents: what actually stays on your servers

Almost every team that says “our agent is self-hosted” is still sending every prompt and every document to someone else’s API. That is not a criticism, it is usually the right call. But you should know which one you bought.

Zakaria El Asri13 min

The sentence that reframes the whole project

Self-hosting the agent and self-hosting the model are two different decisions. One is free. The other starts at 574.87 EUR a month before you process a single request.

In one paragraph

The short answer

"Self-hosted AI agent" describes three different architectures. In the common one, the orchestration runs on your server and the model is still an API call, so your data leaves on every request. In the second, you add a zero data retention agreement and the data leaves but is not kept. In the third, the model runs on your own GPU and nothing leaves at all. Only the third is what most people picture, and it is the only one that costs real money.

The model

An agent has three layers, and you can host them separately

This is the distinction that resolves most arguments about self-hosting, and it is rarely drawn explicitly.

LayerWhat it doesCan you host it?Cost of hosting it
OrchestrationThe loop, the triggers, the integrationsYes, easilyA small VM, tens of euros
ModelThe reasoningYes, with a GPUFrom 574.87 EUR a month
DataYour documents, your databaseYes, you already doAlready in your budget
GPU figure from Scaleway L4 instance pricing, Paris zone, consulted 8 October 2026.

Hosting layers one and three is cheap and almost always worth doing. Hosting layer two is an order of magnitude more expensive and is the one people mean when they say "self-hosted AI".

The expensive misunderstanding

Self-hosted n8n calling an API is not a self-hosted agent

The most common build we are asked to review looks like this: n8n installed on the company's own server, workflows in their own database, logs in their own infrastructure, and an AI Agent node pointing at a hosted model API.

Everything about that is reasonable. The orchestration genuinely is self-hosted, which is a real security and compliance benefit, covered in our evaluation of n8n for AI agents.

But on every single turn of the agent loop, the prompt leaves. And the prompt contains whatever the agent just read: the contract, the patient record, the CV, the client email. The tool result leaves with it, because it is appended to the conversation.

Teams discover this at the worst possible moment, which is during a security review or a client questionnaire, after the system is live. The fix at that point is architectural, not a setting.

Be precise

What leaves, and what does not

ArchitecturePrompts leave?Documents leave?Logs leave?Monthly floor
Hosted platform, hosted modelYesYesYesPlatform subscription
Self-hosted orchestration, hosted modelYesYesNoA small VM
Self-hosted orchestration, hosted model, ZDR agreementYes, not retainedYes, not retainedNoA small VM
Self-hosted orchestration, self-hosted modelNoNoNo574.87 EUR
Monthly floor for the last row from Scaleway L4-1-24G, 0.79 EUR per hour, listed at roughly 574.87 EUR per month running continuously, consulted 8 October 2026.

Note the third row. It is the one most teams should be in, and the one almost nobody knows exists.

The lever nobody pulls

The agreement nobody thinks to ask for

Anthropic's commercial data handling page, updated 1 July 2026, states that for API users inputs and outputs are automatically deleted from the backend within 30 days of receipt or generation. It lists the exceptions plainly: services with longer retention under your own control such as the Files API, legal obligations, usage policy enforcement, and cases where you and they have agreed otherwise.

That last exception cuts both ways, because one of the things you can agree is a zero data retention agreement. It is named in their own documentation as an available arrangement. Most teams building on the API have never asked for it.

For a legal, HR or healthcare workload this changes the conversation materially. Thirty-day retention is often acceptable with the right paperwork; zero retention removes the objection almost entirely, and it costs nothing but a conversation with a sales team.

It does not remove the transfer question, the processor agreement, or your own obligations. The full set of what to get in writing, provider by provider, is in GDPR and AI agents.

The real number

What a private model actually costs

This is where self-hosting stops being free. A model needs a GPU, and a GPU needs to be running whether or not anyone is using it.

InstanceGPUsRAMPer hourPer month, continuous
L4-1-24G148 GB0.79 EUR~574.87 EUR
L4-2-24G296 GB1.58 EUR~1,149.75 EUR
L4-4-24G4192 GB3.15 EUR~2,299.50 EUR
L4-8-24G8384 GB6.30 EUR~4,599 EUR
Scaleway GPU instance pricing, PAR-1 zone, consulted 8 October 2026. Monthly figures are the provider's own stated approximations for continuous running.

Read the first row carefully, because it is the floor and not a typical configuration. 574.87 EUR a month buys one 24 GB card, which runs small or quantised models comfortably and the current frontier reasoning models not at all. Matching the quality your team expects usually means the third or fourth row.

And that is before the parts nobody budgets: someone to keep the inference server running, the model updates, the capacity planning for concurrent requests, and the fact that an idle GPU bills exactly like a busy one. Serving is usually done with vLLM or Ollama, both of which are free and neither of which runs itself.

Compare against the API side honestly. For the volumes a mid-size company generates, token costs typically land well below the GPU floor, and prompt caching lowers them further. The arithmetic is laid out in Claude API pricing for business.

When to pay it

When full self-hosting is justified

There are cases where the GPU bill is simply the cost of being allowed to operate.

SituationWhy hosted will not do
Health data under HDS certificationThe hosting itself is certified and audited; an API call leaves that perimeter
Legal professional secrecyThe duty is personal to the lawyer and does not transfer to a processor by contract alone
Defence, sovereignty, classifiedThe requirement is usually network isolation, not a contract clause
Client contract forbidding third-party processingWritten into the deal, not negotiable by you
Very high, steady volumeAt enough tokens per month the GPU is genuinely cheaper

The first two are the ones we see in practice. The health case has its own architecture, covered in AI agents and health data under HDS. The legal case is subtler than most vendors admit and we wrote it up in AI in law firms and professional secrecy.

Note what is not on that list: "our data is sensitive", "we are a European company", and "the board is nervous about AI". Those are real concerns and they are usually answered by the third row of the earlier table, not by buying GPUs.

The uncomfortable part

When it is not justified, which is most of the time

Most requests we receive for a fully self-hosted agent are driven by a feeling rather than a constraint. That feeling is legitimate and the answer to it is usually cheaper than a GPU.

Three questions settle it. Is there a written rule, in a contract, a certification or a regulation, that forbids a processor? Would a zero data retention agreement satisfy whoever is worried? Does your monthly token volume exceed the GPU floor? If all three answers are no, you are about to pay 575 EUR a month and accept a weaker model to solve a problem you do not have.

Saying this costs us work, because the fully self-hosted build is the larger project. We say it because the alternative is a system that is expensive, slower, and abandoned in eighteen months.

Our position

What we actually deploy

For most clients in regulated sectors, the architecture that survives contact with a security review looks like this:

  • Orchestration self-hosted. n8n on their infrastructure, or our own code, with the workflow database and the logs in their network.
  • Data self-hosted. Documents, the fact store and the audit trail never leave. This is also what makes decisions explainable afterwards.
  • Model over API, with the paperwork done. Processor agreement, documented retention, and a zero data retention agreement where the sector calls for it.
  • A redaction step before the call where it is feasible, so identifiers never reach the prompt in the first place.
  • A private model only where a written rule demands it, budgeted honestly, with the quality tradeoff stated before anyone signs.

That is what we build as custom AI agents, and how we work as an n8n agency. The production checklist that goes with it is in deploying an AI agent to production, and the security surface in AI agent security.

FAQ

Questions about self-hosted AI agents

Usually not. In practice "self-hosted agent" almost always means the orchestration runs on your infrastructure, typically self-hosted n8n or your own code, while the model is still called over an API. Your prompts and your documents still leave your network on every request. Self-hosting the model is a separate, much more expensive decision.

Related guides

Read next

Sources

Links verified at publication. Regulatory texts change — always defer to the official source.

Let's talk about your project

A question, a project, an idea? We respond within 24h. Free audit, no commitment.

Contact details