The real cost of on-premise AI in a company: seven lines missing from the first invoice
The price of the server is the easy part.
A server with an 80 GB NVIDIA H100 GPU is quoted in the Spanish market at between €30,000 and €55,000, depending on CPU, RAM and storage. An L40S with 48 GB brings the price down to €12,000-18,000. Add an industrial-grade UPS, a rack, extra cooling and a 10 Gbps internal network to serve the model, and the first invoice for running a small model (7B-14B parameters) in your own offices comes to around €55,000-95,000. That is the number in the supplier's PDF.
It is a third of the real cost of the first year. Sometimes less.
The rest is seven lines that never appear in that first offer: qualified staff, MLOps, prompt revalidation when the version changes, security, backups, continuity and updates. Each one can be documented with public market ranges. Each one adds between €40,000 and €250,000 to the annual bill, depending on the scenario. Add it all up and the first-year TCO of a local SLM in a mid-sized Spanish company sits between €180,000 and €450,000. The second year comes down a little, but not as much as you might expect.
This post breaks down the seven lines, with figures, and sets them against the cost of processing the same monthly volume through an API under a zero data retention contract. At the end there are three scenarios in which local does pay off. Outside those three, it doesn't.
1. The visible invoice: hardware, energy and network
I'll start with the number that does appear in the quote, because you need it clear before you size the rest.
GPU hardware. A node with one H100 SXM comes to around €40,000 amortised over 4 years. A node with two L40S cards for parallel inference, about €30,000. If your workload is document-based (RAG, summaries, extraction) rather than training, two L40S cards give you more room than a single H100. Add a Xeon or EPYC CPU, 256-512 GB of RAM and 4-8 TB of NVMe storage in RAID. Realistic price for a complete node: €45,000-65,000.
Energy. An H100 draws 700W under sustained load. Once you add CPU, memory, fans and data centre overhead, that rises to 1,100-1,400W per node. With the average Spanish industrial electricity price at around €0.15-0.20/kWh depending on time band and contract, a node running 24/7 spends between €1,500 and €2,400 a year on electricity alone. Double it if you have active redundancy.
Network and data centre. A cabinet in your own data centre with proper cooling, or the equivalent rental in a Spanish data centre (Interxion, Equinix, Global Switch), costs between €4,000 and €12,000 a year depending on location and contracted power. If the data centre is your own, the line disappears into overheads, but it is still there.
Base software and licences. Operating system (RHEL or Ubuntu Pro with support), CUDA driver, inference framework (vLLM, TGI, Triton), orchestrator (Kubernetes or Nomad). Anywhere from free to €8,000 a year if you buy enterprise support from Red Hat or SUSE.
Total visible first year, mid scenario: €65,000-110,000. This is the figure that usually appears in the supplier's first three slides.
2. MLOps staff: the line that multiplies the bill
This is where the trouble starts. A local model is not something you install once and forget. It needs a person or a team who can deploy, monitor, update, measure latency and quality, and respond when something breaks at 11pm on a Tuesday.
Salary ranges for MLOps profiles in Spain, according to the latest Michael Page reports and the Hays Salary Guide:
| Profile | Gross annual salary (Madrid/Barcelona) | Freelance rate/hour |
|---|---|---|
| Junior MLOps Engineer (2-3 years) | €35,000-45,000 | €45-65 |
| Senior MLOps Engineer (5+ years) | €55,000-75,000 | €75-110 |
| ML Platform Engineer / Lead | €70,000-95,000 | €100-150 |
| SRE with GPU specialisation | €60,000-85,000 | €85-130 |
In plain terms: hiring a senior MLOps engineer on the payroll, at total employer cost (roughly 1.32x gross salary), comes to between €73,000 and €99,000 a year. If you go freelance because you can't find the profile or don't want to add headcount, a part-time contract (20 hours a week) runs to around €80,000-115,000 a year.
One MLOps engineer on their own is not enough. At a minimum you need:
- A senior MLOps or platform engineer who owns deployment and updates.
- An on-call SRE, shared with the rest of the infrastructure, who knows GPUs.
- A data engineer to maintain the ingestion pipeline into the model (RAG, embeddings, reindexing).
Staff cost in year 1, minimum scenario with one full-time senior MLOps engineer plus time from the existing SRE: €85,000-110,000. Realistic scenario with a dedicated team of two: €150,000-200,000.
Rule: if your company doesn't already have a platform team that can run GPU workloads, this is the line that decides whether the project lives or dies. Training courses won't replace it.
3. Prompt revalidation when the model changes
Nobody sees this cost until month 8, and it catches everyone out.
Open models get updated. Llama 3 becomes Llama 3.1 and then 4. Mistral 7B becomes Mistral Small 3. Qwen ships two versions a year. Every time you upgrade the base model for better quality, longer context or lower compute cost, the prompts you had tuned stop behaving as they did.
Symptom: the prompt that correctly extracted 94% of the amounts on an invoice now extracts 81%, or it extracts 96% but returns the result in a different format that breaks the system downstream.
Fix: a full revalidation cycle. And that cycle costs money.
A typical revalidation cycle for a company with 12 use cases (extraction, classification, summarisation, customer replies, proposal drafting and so on) involves:
- Evaluation suite. A labelled dataset of 200-500 examples per use case. If you don't have one, it has to be built. One-off cost of €8,000-25,000 in human hours. Annual maintenance: 20-30% of the initial cost.
- Eval framework. LangSmith, Promptfoo, DeepEval or something built in-house. Anywhere from free to €12,000 a year in licences.
- Prompt engineering hours. Each use case under review takes between 4 and 20 hours from someone who knows what they are doing (ML engineer or prompt engineer). At an internal cost of €80-120/hour, with 12 cases, each revalidation cycle costs between €6,000 and €25,000.
- Realistic frequency. At least two full cycles a year if you want the improvements coming out of the open ecosystem. If you stay anchored to the model you deployed on day one, you fall 40% behind the state of the art in quality within 18 months.
Annual prompt maintenance cost for a local deployment with 12 use cases: €20,000-60,000.
Critical note: if you work against a closed API, this cost also exists when the provider updates the model underneath you. The difference is that with an API it is optional (you can pin a versioned snapshot for 6-12 months), whereas on premise it is mandatory if you want to stay close to the state of the art.
4. Security, backups and model isolation
When the pitch for bringing the model on premise is "the data is sensitive and never leaves the company", consistency means building security that matches the pitch. If the model processes clinical records, payroll or public tender contracts, it cannot sit on the same flat network as the ERP and the printer in reception.
Network segmentation. Dedicated VLAN, internal firewall, role-based access rules. Initial set-up: €6,000-15,000 in network consultancy. Annual maintenance: €3,000-8,000.
Encryption at rest and in transit. The embedding volumes and the vector index contain fragments of the original content and are as sensitive as the source documents. LUKS encryption or equivalent, key management with HashiCorp Vault or a similar tool. Annual cost between €4,000 and €15,000 depending on the tool.
Backups and disaster recovery. The vector index of a RAG system with 200,000 documents takes up 40-120 GB and needs 8-30 hours to reindex from scratch. If it gets corrupted and there is no backup, your system is out of action for a full working day. Daily incremental backup, weekly snapshot, quarterly restore test. Annual storage and tooling cost: €5,000-18,000.
Audit log. A record of which user asked what, when, and what the model answered. Mandatory for high-risk use cases under the EU AI Act (recruitment, healthcare, essential services) and strongly advisable for everything else. Retention of at least 24 months. Annual cost between €6,000 and €20,000 depending on volume and tool.
Annual penetration test. A local model service with access to internal data is a new attack surface. A dedicated pentest focused on prompt injection, exfiltration via embeddings and context abuse costs around €8,000-18,000 in the Spanish market, depending on scope (indicative pricing from firms such as S21sec or Deloitte Risk).
Reasonable annual security cost for a local deployment with sensitive data: €32,000-74,000.
If the pitch can't absorb this line, it was never a real pitch.
5. Updates and continuity: the line nobody budgets for
A local model ages fast. I don't mean only the base model. I mean the whole stack.
CUDA driver and operating system. Updates every 3-6 months. Each one means checking that inference keeps the same latency and accuracy. Cost in hours: €2,000-6,000 per cycle.
Inference framework. vLLM ships a minor release every 4-6 weeks and a major one every 6 months. TGI and Triton are much the same. Any major release can change the internal API and break your wrapper.
Base model. At least one migration a year if you want to stay close to the state of the art. Each migration is a 3-8 week project, including prompt revalidation (covered in point 3) and latency and throughput testing under real load.
Hardware support contract. Extended warranty and GPU replacement within 24-48h. Without it, a failed GPU can leave you out of service for 3-8 weeks while you wait for repair or replacement. Annual cost: 4-8% of the hardware value. On a €50,000 node, between €2,000 and €4,000 a year.
Active redundancy. If the system is critical (internal users depend on the model to do their jobs), you need at least two nodes, with a load balancer and failover. That doubles the hardware, the energy and the support licences. Multiply by 1.8 (not 2, because part of the data centre and the team is shared).
Annual maintenance and update cost, excluding the human team (already in line 2): €12,000-35,000. With redundancy, base hardware goes up by 80%.
6. API with zero data retention: the comparison almost nobody gets right
The headline argument for going local is "I can't send my data to an American API". Three years ago that argument carried real weight. In 2026, every serious provider offers enterprise contracts with Zero Data Retention, EU data residency and, in some cases, residency in Spain.
OpenAI Enterprise. Contract with Zero Data Retention, optional EU residency, SOC 2 Type II and ISO 27001 compliance. Data is not used for training and is not kept beyond the inference window.
Anthropic Claude for Enterprise. Via AWS Bedrock with residency in Frankfurt or Ireland. Enterprise contract, Zero Data Retention, HIPAA and SOC 2 compliance.
Azure OpenAI Service. Residency in Spain Central (a new region since 2024) or Sweden. Microsoft enterprise contract, integrates with Purview for data governance.
Mistral La Plateforme (enterprise). French company, servers in the EU, Zero Retention contract, a European model by default.
All four will sign a DPA (Data Processing Agreement) aligned with the GDPR, with a sub-processor annex. A Spanish data protection consultant can validate the contract in 2-4 hours.
Now the numbers. A realistic case: a mid-sized company processing 2 million pages a year through an extraction and summarisation pipeline (invoices, contracts, sales emails). Estimated token volume: 3,000 million input tokens and 400 million output tokens a year.
With a model such as GPT-4o mini, Claude Haiku 3.5 or Mistral Small on an enterprise API, the combined price ranges from €0.15 to €0.60 per million input tokens and from €0.60 to €3 per million output tokens, depending on provider and negotiated volume.
Annual API cost for that volume:
- Cheap scenario (Mistral Small or GPT-4o mini with a volume discount): 3,000 M × 0.15 + 400 M × 0.60 = €690 a year (yes, six hundred and ninety).
- Mid scenario (Claude Haiku 3.5 enterprise): 3,000 M × 0.80 + 400 M × 4 = €4,000 a year.
- Premium scenario (GPT-4o or Claude Sonnet 4.5 on 100% of the volume): 3,000 M × 2.50 + 400 M × 10 = €11,500 a year.
On top of the API, add the human cost of maintaining the pipeline (it exists here too, but it is smaller): a part-time data engineer (0.3 FTE) for monitoring, throttling, error handling and fallback. Around €18,000-25,000 a year.
Total API with an enterprise contract and part-time staff for 2 M pages/year: between €19,000 and €37,000 a year.
Total local for the same volume, realistic scenario, year 1: between €210,000 and €380,000.
The difference is not marginal. It is an order of magnitude, except at very high volumes.
7. When a local SLM does pay off: three scenarios
I am not saying local never makes sense. I am saying that most companies considering it are in none of the three scenarios where it pays off. Here they are:
Scenario A: massive volume with predictable loads. You process more than 20 million pages a year, or more than 40,000 million input tokens, with sustained load around the clock. Beyond that point, per-token inference through an API clearly costs more than the local TCO amortised over 4 years. Typical profile: a large insurer reading claims, a legal platform analysing case law at scale, a media group processing its archive. Few mid-sized companies are in that position.
Scenario B: regulated data that literally cannot leave a physical perimeter. Specific cases: information classified under official secrecy, military data, some biomedical research funded by consortia with specific physical isolation clauses. Note: the GDPR and ordinary health data do not require this. An enterprise contract with EU residency and encryption in transit and at rest covers almost the entire Spanish health sector, provided there is a signed DPA.
Scenario C: latency below 50 ms with continuous traffic from a single building. Real-time applications on the factory floor, robotics, the odd point-of-sale assistant with a very tight SLA. Here the network round trip to a public API, even one hosted in a European data centre, doesn't fit. Local inference solves it. It accounts for a very small share of generative AI use cases in companies.
Outside these three, the local TCO does not pay off in 2026 for a mid-sized Spanish company.
Common mistakes when deciding to buy a GPU server
Mistake 1: Comparing the hardware price against the unit price of the API. Symptom: the supplier's spreadsheet only shows server versus cost per token. Fix: demand a 3-year TCO with the seven lines from this post broken down and signed by the supplier.
Mistake 2: Assuming "on premise = GDPR compliant". Symptom: the salesperson argues that bringing the model into your offices settles the data protection question. Fix: GDPR compliance comes from a legal basis, a DPA, encryption, data minimisation and traceability, not from where the server physically sits. A badly configured local deployment breaches it as much as, or more than, a well-contracted enterprise API.
Mistake 3: Not budgeting for the human team. Symptom: the proposal includes hardware, installation and "3 days of training for the current team". Fix: budget for at least one full-time senior MLOps engineer from month 1 and a shared SRE who knows GPUs. If there is no budget for that, there is no budget for local.
Mistake 4: Choosing the biggest model "just in case". Symptom: the supplier proposes Llama 3.3 70B because "it has more quality" when the use case is extracting structured data from invoices. Fix: start with a 7B-14B model, measure accuracy with a real evaluation suite and move up in size only if the numbers justify it. Doubling the model size quadruples the GPU you need.
Mistake 5: Ignoring the cost of migrating back. Symptom: nobody asks what happens when you decide to switch off the local system in 24 months because API prices have fallen by another 60%. Fix: insist on a decoupled architecture from the design stage (an abstraction layer over the model provider, versioned prompts, a portable evaluation suite) so you can migrate without rewriting the pipeline. Extra upfront cost: 4-8 weeks of work. Potential saving: 6 months of re-engineering on the day you want to switch.
Frequently asked questions
How much does it really cost to run a small local model in a Spanish company?
In a mid-sized company (100-500 employees) with a document-based use case and no existing MLOps infrastructure, the first-year TCO sits between €180,000 and €450,000, with €65,000-110,000 in visible hardware and the rest in staff, security, maintenance and revalidation. The second year comes down by 20-30% once set-up and the initial investment in the evaluation suite drop out. The ongoing figure settles at around €140,000-320,000 a year.
What hidden costs come with running your own LLM on a server?
The seven that never appear on the first invoice: MLOps staff (€85,000-200,000/year), prompt revalidation (€20,000-60,000), security and isolation (€32,000-74,000), backups and disaster recovery (€5,000-18,000), stack and model updates (€12,000-35,000), hardware support and extended warranty (€2,000-4,000 per node) and energy and data centre (€5,000-14,000 depending on redundancy). No hardware supplier budgets for the first five because they are not its business.
Is a local SLM or an OpenAI API cheaper for processing documents?
For volumes below 20 million pages a year (95% of mid-sized companies), an API with an enterprise contract and Zero Data Retention is 8 to 20 times cheaper than local, on a full TCO basis. Above 20 million pages a year with sustained load, the break-even point gets closer. Above 50 million with 24/7 load, local usually wins when amortised over 4 years.
What does a company need to keep an AI model running on premise?
Four minimum blocks: people (at least one dedicated senior MLOps engineer and a shared SRE who knows GPUs), technical stack (inference framework, orchestrator, monitoring, secrets management, encryption), processes (versioned evaluation suite, prompt revalidation at least twice a year, a disaster recovery plan tested every quarter) and contracts (24/7 hardware support, extended warranty, annual pentest, internal DPA with the user departments). If any of the four is missing, the deployment works for the first three months and degrades silently from month six onwards.
Can I comply with the GDPR using an API instead of a local server?
Yes, and in most cases with less effort. The four big enterprise providers (OpenAI, Anthropic via Bedrock, Azure OpenAI, Mistral) offer contracts with Zero Data Retention, EU residency and a DPA you can sign. A Spanish data protection consultant validates the contract in 2-4 hours. The specific cases where regulation does require local: information classified under official secrecy, military data and some biomedical research consortia with specific clauses. Healthcare, banking, insurance and ordinary public administration all fit within enterprise API contracts with EU residency.
How long does a GPU server bought in 2026 take to pay for itself?
It depends on volume. With the average load of a mid-sized company (2 million pages a year), it never pays for itself. The API with an enterprise contract always wins. With high load (20-40 million pages a year), estimated payback is 3.5-4 years, just as the hardware reaches end of life and needs replacing. With massive, sustained 24/7 load (50+ million pages/year), payback comes in 24-30 months. The hardware supplier rarely includes this calculation in its quote.
Closing: the number to ask for before you sign
The PDF proposal with the price of the GPU server is not bad faith on the supplier's part. It is their business model. They sell hardware, and their quote covers what they supply. The real cost of the project sits in the seven lines they don't sell, which land on your operating budget the following year.
Before signing any on-premise AI proposal, ask the supplier for a 3-year TCO with the seven lines from this post broken down: hardware, energy, staff, revalidation, security, backups and updates. If they can't or won't sign it, they are not the right supplier. If they sign it and it still comes out ahead of an enterprise API contract with Zero Data Retention for the same volume and use case, go ahead. In most cases it won't.
If you need to review a specific GPU server proposal before signing, or compare the local TCO with an enterprise API architecture with EU residency for your real volume, you can book a free 30-minute session to review your case. The session ends with a written calculation of the seven lines applied to your figures and a practical recommendation: sign, renegotiate or change the architecture.
Related reading
- Mandatory AI literacy: what Article 4 of the EU AI Act requires
- Enterprise API contracts with Zero Data Retention: what to check in the DPA
- Prompt evaluation suite: how to build your own in 3 weeks
Sources
- EU Artificial Intelligence Regulation (Regulation (EU) 2024/1689, EU AI Act)
- Michael Page, Estudio de Remuneración Tecnología 2026 (Spanish technology salary survey)
- Hays Salary Guide Spain 2026
- INCIBE (Spain's national cybersecurity institute), cybersecurity services for companies
- Public documentation from OpenAI Enterprise, Anthropic Enterprise, Azure OpenAI Service and Mistral La Plateforme on Zero Data Retention and EU residency
