Saltar al contenido
APFerrer

AI model gateway: how to design an enterprise architecture without vendor lock-in

APFerrerSeptember 30, 202618 min
Lead

A model gateway routes each call to the right provider and lets you switch vendor without rebuilding the system. Lock-in doesn't come from using an API. It comes from gluing your architecture to a single provider.

AI model gateway: how to design an enterprise architecture without vendor lock-in

"Architecture is the sum of the decisions that are expensive to reverse."

In 2024, OpenAI raised the price of GPT-4 Turbo by 50% over the previous version in some usage tiers, and during 2025 it changed the header structure for function calling three times. Anthropic renamed the Claude 3 family to Claude 3.5, then to Claude 4, and along the way moved the system parameter from a string to a list of typed blocks. Google migrated from the PaLM API to the Gemini API with a breaking contract change. Mistral went from an OpenAI-style API to adding its own JSON schema for tool use. In eighteen months, any team that wrote code directly against an official SDK has had to touch its integration layer between four and six times.

The usual reading in architecture meetings is: "this happens because the market is young". The correct reading is different. The market will keep changing like this for at least the next five years, because the marginal cost of a token has fallen 40% a year since 2023 according to the LLM Price Check index, and every drop forces providers to redesign tiers, limits and features so as not to cannibalise their high end. If your architecture absorbs that noise with every release, you don't have a provider problem. You have a design problem.

Lock-in doesn't live in the API call. It lives in the decisions about prompts, parsing, tokens and caching that you tie to one vendor's quirks. A gateway layer makes those decisions neutral. That's all.

Gateway, router and adapter: three pieces, three responsibilities

The first confusion I see in every architecture review is treating these three words as synonyms. They aren't. Each one solves a different problem, and collapsing them into a single module is why so many home-made gateways end up as an if provider == "openai" inside a helper.

Adapter. Translates your application's internal contract into a specific provider's contract. An Anthropic adapter knows that system goes outside the messages array. An OpenAI adapter knows that tools is a list of objects with type: "function". An adapter is dumb code with no business logic: it takes a neutral message and returns a neutral response. One adapter per provider. No exceptions.

Router. Decides which adapter each request goes to, following rules: task type, budget, target latency, availability. The router knows nothing about how Anthropic serialises tool_use. It knows that the "classify intent" task goes to a cheap model and the "draft legal proposal" task goes to a large one. The router's logic is your company's policy, not a vendor's.

Gateway. The network layer that wraps the router and adapters and adds what makes all this workable in production: centralised authentication, per-team rate limiting, semantic caching, observability, auditing, retries and fallback. A gateway is deployed as an independent service and exposes an internal HTTP API. Your applications don't talk to OpenAI. They talk to the gateway.

Translation: the adapter is a translator, the router is a referee, the gateway is customs. Mixing up the three layers is the mistake that makes "switching provider" take three weeks instead of three hours.

A mid-sized team with fifteen developers and four products consuming LLMs saves a full release per quarter just by separating these three pieces properly. This isn't theory. It's what any team measures after moving from "official SDK in every service" to "HTTP client calling an internal gateway".

The neutral message contract: the abstraction everything else depends on

A gateway without a neutral contract is a proxy. And a proxy doesn't free you from lock-in. It only moves the problem one layer along. The piece that really matters is the internal message schema your application produces before the request reaches the gateway, which the gateway consumes without depending on any vendor.

A minimum viable contract has seven fields:

  • role: system, user, assistant, tool. Four values. Zero variations per provider.
  • content: a list of typed blocks (text, image, tool_use, tool_result). Not a plain string. Plain strings are the first trap.
  • tools: a list of tool definitions in strict JSON Schema. No embedded type: "function", no parameters with proprietary extensions.
  • params: temperature, max_tokens, stop sequences, response_format. A shared vocabulary, not the one from whichever SDK you read last.
  • metadata: task_id, tenant_id, user_id, cost_budget. What the router needs to decide.
  • cache_key: a deterministic hash of the request's semantic content. Without it, the cache can't live in the gateway.
  • schema_version: an integer you bump when you change the contract. Without it you can't migrate without breaking older clients.

A real example. A payroll fintech with six internal services uses LLMs to classify tickets, extract data from payslip PDFs, draft replies to incidents and audit changes to employee records. Before the gateway, every service imported the OpenAI SDK and built the messages object by hand. When the team decided to try Claude for PDF extraction (it performs better on long documents), they had to rewrite the helper in every service. Three sprints, two production regressions and one billing incident caused by miscalculating the new token scheme.

After the gateway, each service sends an object with those seven fields to /v1/complete, and the gateway hands it to the right adapter. Switching provider for the PDF flow is one line in the router's rule table: task_type = "pdf_extract" -> provider = "anthropic". Zero changes in the consuming service. Zero regressions in the other flows.

Rule: if your internal contract looks like a particular provider's schema, you don't have an internal contract. You have that provider in disguise. The test is to explain the schema to someone without telling them which vendor you use. If they guess the vendor, you've failed.

Typed blocks in detail

The quietest mistake when designing the contract is accepting content as a string. It works for simple chat. It breaks as soon as vision, tool use, thinking blocks, citations or any multimodal feature comes in. Every serious provider already supports at least three of those five.

A typed block has the shape { "type": "...", ...fields }. A text block is { "type": "text", "text": "..." }. An image block is { "type": "image", "source": { "media_type": "...", "data": "..." } }. A tool call block is { "type": "tool_use", "id": "...", "name": "...", "input": {...} }. The adapter flattens or expands them to match what the vendor expects.

Providers that still accept a plain string in content when there's only text do so for backwards compatibility. None of them recommends that format for new code. Designing your contract around a list of blocks from day one costs almost nothing. Redesigning it once you have twenty flows in production is a migration measured in months.

Routing by cost, latency and availability: policy, not whim

A serious router decides with three measurable variables and a versioned rule table. No "let's try a few and see which works best". No "the CTO prefers Anthropic". Decisions are made with data and audited.

Cost per token. Every provider publishes its rate per million input and output tokens. In September 2025 the range runs from $0.15 per million input tokens on small models (Haiku 3.5, GPT-4o mini, Gemini Flash) to $15 per million on large models (Opus 4, GPT-4 Turbo, Gemini Pro 1.5). A hundredfold gap between the cheapest and the most expensive. If your router sends every request to the large model "just in case", your monthly bill is twenty times what it should be.

Observed latency. Not the figure the provider quotes in its marketing. The one you measure, in your region, at your volume. A model whose API responds in 1200 ms at the median is useless for a conversational chatbot that needs the first token in under 400 ms. Measure the 95th percentile over the last seven days, not the average, because the median hides the long tails that ruin the experience.

Availability. The share of 2xx responses over total requests. A provider with 99.5% availability in production is down for ~3.6 hours a month. If that takes out a critical flow, you need automatic fallback to another vendor. No exceptions.

In the minimal case, the router's rule table looks like this:

Task type Primary provider Fallback Max latency P95 Max cost per call
chat_customer openai/gpt-4o-mini anthropic/haiku-3.5 800 ms €0.002
pdf_extract anthropic/sonnet-4 openai/gpt-4o 12000 ms €0.05
classify_intent mistral/small openai/gpt-4o-mini 400 ms €0.001
generate_report anthropic/opus-4 openai/gpt-4-turbo 30000 ms €0.50
translate_es_en google/gemini-flash openai/gpt-4o-mini 600 ms €0.001

That table lives in configuration, not in code. It changes with a two-line pull request, deploys in minutes and leaves an audit trail in git. Adding a new local model (say, Llama 3.3 served on your own infrastructure) means adding a row. Nothing more.

A hard budget per tenant

On a platform with several clients or internal teams, the router should read a monthly budget per tenant from the request metadata. Once a tenant has spent 90% of its quota, the router quietly drops down to the cheapest model in the same functional group. Past 100%, it returns a 402 with a clear message. In 2026 this isn't optional. Without a hard budget limit, a runaway loop (a badly written service calling the LLM inside a while True) drains the card before the billing alert reaches the CFO's inbox.

Tokens, caching and observability: accounting that can't live in the client

The second silent lock-in sits in the accounting. Each provider counts tokens with a different tokenizer: tiktoken for OpenAI, Anthropic's own tokenizer (available through its token counting API), SentencePiece for Gemini, a different SentencePiece model for Mistral. If your application estimates tokens on the client before calling, you're importing the vendor's tokenizer and tying yourself to it.

Fix: token counting lives in the gateway. Always. The client sends the neutral message and the gateway returns, alongside the response, a normalised usage block: input_tokens, output_tokens, cached_tokens, cost_eur. The client calculates nothing. The gateway logs every call in an audit table with tenant_id, task_id, provider, model, usage, latency_ms, cache_hit.

Semantic caching, not textual

A naive cache stores hash(prompt) -> response. It works for requests that are identical byte for byte. In a real product, two users ask the same thing with different spacing and punctuation, and the cache misses. A semantic cache hashes a canonical form of the message, after normalising whitespace, lowercasing and sorting JSON fields in a stable order. In RAG flows, it also hashes the ordered list of retrieved document IDs.

With semantic caching on a platform handling 200,000 requests a day, the typical hit rate is between 15% and 35% depending on the type of flow. At the top of that range, the direct saving on the bill runs to thousands of euros a month. The hit rate climbs to 60% in flows with long prompts and short variables (typical of document assistants), because the cached response depends on the content, not on the vendor.

Critical note: the cache has to live in the gateway, not in the client. If it lives in the client, each application keeps its own cache, you lose deduplication across products, and switching provider throws away caches that took weeks to warm up.

Vendor-agnostic observability

LLM observability isn't the same as observability for an HTTP microservice. You need five specific metrics:

  • Cost per task type. Sum of cost_eur per hour, grouped by task_id. Catches leaks.
  • Latency P50/P95/P99 per provider. Used to tune the router's rule table.
  • Cache hit rate per task type. If a flow drops below 10%, check the cache key.
  • Token efficiency ratio. output_tokens / input_tokens. A low ratio (0.1) points to prompts padded with unnecessary context. A high ratio (5+) points to open-ended generation with no clear limit.
  • Error rate per provider. Split by status code: 429 (rate limit), 500 (vendor), 400 (contract), 401 (auth). Each one needs a different response.

All five are exported via OpenTelemetry from the gateway. Any observability stack (Grafana, Datadog, Honeycomb, the open source tool of the moment) consumes them the same way. Nothing in your application knows they exist. Switching provider changes the data, not the instruments.

Migrating a production flow step by step, with the lights on

How you migrate a flow between two providers in production is the best test of a gateway design. If the migration is an overnight event with a deployment freeze, the design is wrong. If it happens during working hours with one-click rollback, it's right.

Scenario: an administrative consultancy runs an internal assistant that answers staff questions about employment law, internal procedures and payroll status. The current flow uses OpenAI GPT-4o for every answer. Volume: 8,000 requests a day, monthly cost ~€900. The team wants to try Anthropic Sonnet 4 because an offline evaluation shows better accuracy on answers that cite collective bargaining agreements.

Step 1: shadowing. The router is set to send 100% of traffic to OpenAI (the response the user sees) and, in parallel and asynchronously, replay each request against Anthropic. Anthropic's response is never served to the user. It's stored in a shadow_responses table next to OpenAI's. Duration: two weeks.

Step 2: evaluation. A daily job compares both responses using an evaluator (another LLM with a strict prompt, or a human reviewing a random sample of 200 cases). It measures factual accuracy, adherence to the corporate tone, latency and cost. If Anthropic wins on three of the four metrics, the change is approved.

Step 3: canary. The router moves from 100% OpenAI to 90% OpenAI / 10% Anthropic. Canary users are picked at random, identified by a hash of the user_id. Duration: one week. If the quality metrics hold and there are no incidents, the split goes to 50/50.

Step 4: full rollout. 100% Anthropic. OpenAI stays as the fallback in the rule table. If Anthropic fails or goes over budget, the router switches to OpenAI automatically.

Step 5: clean-up. After a stable month, shadowing is switched off and the comparison tables are archived. Prompts are reviewed to take advantage of Anthropic-specific features (for example, prompt caching on long system blocks, which cuts the cost by another 30%).

Total duration: five weeks. Zero downtime. Zero changes to the consuming application's code. Every change lives in the router's rule table, in versioned configuration. Rollback at any point is a git revert and a deployment.

This sequence isn't a theoretical ideal. It's what any serious platform team with a well-built gateway does. The only precondition is having taken the three previous points seriously: a neutral contract, rule-based routing and centralised accounting.

Three typical mistakes that let lock-in back in through the back door

A well-designed gateway can be undermined in two weeks if the team doesn't watch for these three patterns. I see them in every review.

Mistake 1: prompts that rely on vendor-specific features

Symptom: the system prompt contains instructions aimed at one provider. Phrases such as "use your thinking block before answering", "reply in structured JSON using your response_format" or "call the tool with Anthropic's XML format". When the router sends that prompt to another vendor, the output degrades because the other vendor doesn't understand the instruction.

Fix: write prompts in vendor-neutral language. "Think before you answer" rather than "use thinking". "Return the answer as a JSON object with these fields" rather than "use response_format json_object". Vendor-specific features are switched on in the adapter, not in the prompt. If the vendor supports native thinking, the adapter enables it from a params.enable_reasoning = true flag. If it doesn't, the generic prompt still works.

Mistake 2: fragile response parsing

Symptom: the consuming code reads response.choices[0].message.content or response.content[0].text directly. When you change vendor, that access breaks. The "simple" fix is a try/except in the client. And you're back where you started.

Fix: the gateway always returns a response with the same schema: { "message": { "role": "assistant", "content": [typed blocks] }, "usage": {...}, "provider": "...", "model": "..." }. The client reads response.message.content and nothing else. Parsing each vendor's response format is the job of the output adapter, not the client.

For structured responses (strict JSON), the gateway validates against the JSON Schema the client sent in params.response_schema. If the vendor doesn't support native structured output, the adapter enforces it with a post-parser. The client always gets valid JSON. Zero per-vendor try/except.

Mistake 3: a cache key that includes vendor data

Symptom: the cache key is built as hash(provider + model + prompt). It sounds sensible. It's a disaster. When the router switches vendor for the same task, the cache doesn't find the result and makes a fresh call. Worse, two identical requests routed to different vendors pay twice for the same answer.

Fix: build the cache key from the canonicalised semantic content and the task_id. Nothing else. The vendor stays out of the key. If a cached response exists, it's served to the client without calling any LLM. If not, the router picks a vendor, makes the call and stores the response under that same neutral key. Next time, the cached answer is served whichever vendor the router would have picked.

Legitimate exception: when vendor A's and vendor B's responses differ consistently in tone, format or accuracy, and the user can see that difference. In that case the cache key does include the vendor, but the cost is accepted explicitly, documented in the rule table and reviewed every quarter.

Common mistakes

Mistake 1: The gateway as a dumb proxy. Symptom: the "gateway" is an nginx with URL rewrites pointing at api.openai.com. Fix: if your gateway doesn't translate contracts, observe, cache and route, it isn't a gateway. It's a proxy. Rebuild it as an application layer in your stack's language (Python with FastAPI, Node with Fastify, Go with chi) that covers all four.

Mistake 2: One default model for everything. Symptom: the rule table has a single entry, "any task -> gpt-4o". Fix: measure cost per task_type for two weeks. Find the three tasks that account for 80% of spend. Move those three to the cheapest model that passes your quality evaluation. Typical saving: 60% of the bill with no change to the experience.

Mistake 3: Fallback without criteria. Symptom: when the primary provider fails, the gateway retries against the same one. Three times. With backoff. The cascade ends with a 500 returned to the user after ten seconds. Fix: fallback fires on the first 5xx or the first timeout, and goes to a different vendor with a model of equivalent quality. The policy is defined in the table, not in client code.

Mistake 4: No versioning of the message contract. Symptom: a new field is added to the neutral schema and older clients start failing with "unknown field". Fix: every message carries schema_version. The gateway accepts N and N-1 for at least a quarter. Deprecation of N-2 is announced two weeks in advance in the internal changelog.

Mistake 5: Auditing without PII scrubbing. Symptom: the audit table stores the full prompt with customer data: emails, national ID numbers, medical histories. When the DPO asks for a report, a legal problem turns up. Fix: the gateway scrubs PII before writing to the audit log. Prompts are hashed or masked. Raw responses are kept for no more than 30 days.

Frequently asked questions

How do you avoid lock-in with an AI provider such as OpenAI or Anthropic?

With your own gateway layer that defines a neutral message contract (roles, typed blocks, JSON Schema for tools) and per-provider adapters that translate that contract into each SDK. Consuming applications only talk to the gateway. You change vendor by editing the router's rule table, without touching business code. The three critical points are prompts without proprietary instructions, unified response parsing, and a cache keyed on semantic content that leaves the vendor out.

What is a model gateway and what does it do for a company?

A model gateway is an internal service that centralises every call to LLM providers. It does four concrete jobs: it mediates the contract between your application and any vendor (adapters), decides which model to use based on cost, latency and availability (router), cuts the bill with a shared semantic cache, and audits every call with consistent metrics. Without a gateway, each team wires in the vendor's SDK by hand, and every price or API change forces a migration across the whole organisation.

How do you switch LLM provider without rebuilding the system?

With a five-step migration: shadowing (two weeks replaying requests against the new vendor without serving them to users), offline evaluation with quality metrics, a canary from 10% to 50%, full rollout, and demoting the previous vendor to fallback. If the gateway is well designed, no consuming service changes a line of code. The changes live in the router's rule table, versioned in git.

Which architecture pattern lets you mix OpenAI, Anthropic and local models in the same product?

The gateway + router + adapters pattern. The gateway exposes a single internal API. The router decides by task_type and a cost/latency policy. Each adapter (one per vendor, plus one per local model served with vLLM, Ollama or TGI) translates the neutral contract into the backend's format. A local Llama model sits in the same rule table as a remote GPT-4o. The consuming application can't tell them apart.

How much does a well-built gateway save on the bill?

It depends on the starting point. For teams spending between $5,000 and $100,000 a month on LLMs that don't yet split tasks by model, the typical saving from a gateway with task_type routing is between 40% and 65% of spend. Semantic caching adds another 15%-30% depending on the type of flow. With no visible change to the user experience.

Is it worth building your own gateway or using a third-party one (OpenRouter, LiteLLM, Portkey)?

It depends on size and sector. A team spending less than $5,000 a month on LLMs, with no strict regulatory requirements, can use a third-party gateway and save time. A team with regulated data (healthcare, banking, public sector), high volume or business-specific routing logic should build its own. The reason isn't features. It's data exposure: the gateway sees every one of your prompts in plain text. That isn't something to outsource lightly.

Closing

Lock-in with an LLM provider is an architecture problem, not a vendor problem. Vendors change their terms every quarter. That's their job, and it's what any infrastructure provider does in a market where prices fall 40% a year. Your job as an architect is to make sure those changes never reach your code. Getting there doesn't take a new platform. It takes three properly separated layers, a neutral contract that survives five SDK changes, and the accounting in the place it belongs.

It's not complicated. It's discipline.

If you need to review the current architecture of your AI platform and decide what to move to the gateway first, which models to consolidate and which routing policies cut the bill without hurting the experience, that's exactly what an AI model orchestration session covers. You leave with a map of your current flows, a proposed rule table and the three priority migrations ranked by impact on cost and risk. If you want to start there, book a session.


Related reading

Sources

AF
APFerrer
APFerrer · Consultora en datos y procesos
Author's note

Does it apply to your company? Tell me in 30 minutes and we'll see what fits.

Book 30 min