Most organizations start with one AI provider. A team picks a model, writes against that provider's SDK and ships something useful. Then a cheaper model appears, or a better one for a particular task, or a client asks for their data to stay with a provider they already use. A second provider arrives, then a third, and each brings its own SDK, parameter names, keys and limits.
A model gateway puts one layer between your applications and those providers. Applications ask for a model by name. The gateway sends the request to the right provider, applies policy and records what happened. This article describes how to design that layer, based on one we built for a global consultancy whose teams choose from 45 chat models across six providers.
Why direct integrations stop working
Connecting each workflow straight to a provider is the fastest way to start. It becomes expensive in three ways as the number of workflows and providers grows.
Lock-in at the level of code
Each provider has its own SDK, parameters and model names. Code written for one does not move easily to another when prices, quality or availability change, and the cost of switching grows with every workflow.
No single point of control
Keys end up spread across services. Usage is only visible on each provider's bill. Safety policies are implemented separately in every workflow, if at all.
No visibility into agents
Multi-step agents fail in ways that are hard to reproduce. Without a record of each model call and its inputs, debugging becomes guesswork.
The architecture in five parts
A gateway is more than a proxy. The parts below work together so that choosing a model becomes a configuration decision, and oversight sits in one place.
A provider-neutral interface
Applications call one interface instead of a provider SDK. Libraries such as LangChain provide this, and so do gateways that accept a single request format for every provider. The rule is simple: no workflow imports a provider SDK directly.
A model catalogue held as configuration
Describe each model by provider, model name and a small set of behaviour settings: temperature, a seed where output must be reproducible, and a token limit. Validate those settings before a model is created, so a mistake fails when the configuration is loaded rather than in production. Keep reasoning models in a catalogue of their own, because they take different settings and serve different work.
A gateway in the request path
Every model call passes through the gateway, which routes it to the provider, applies the policies attached to the caller and logs the call. Keep provider keys in the gateway or a secret store, so applications never need to hold them.
Policies attached by reference
Define guardrail policies in the gateway, then attach them to an agent or workflow by name. A policy can change without a redeploy, and the same policy applies everywhere it is attached.
Tracing for every chain
Record the prompt, response, timing and model for every step of a multi-step agent. When an agent misbehaves, the trace shows which step went wrong and what it was given. We used LangSmith for this.
Keep the gateway itself replaceable
A gateway removes lock-in to model providers, but the gateway product can become the next lock-in. The answer is the same pattern one level down: put the gateway behind an interface you own.
Changing a model safely
A gateway makes changing a model cheap. It does not make it safe. Two models given the same prompt can differ in accuracy, format, length, cost and speed, and a prompt tuned for one often needs adjusting for another.
Keep an evaluation set for each use
Collect representative inputs with the outputs you expect, including the awkward cases. Run a candidate model against them and compare the results before it replaces the current one.
Record what each model can do
Context window, tool calling, structured output and image input differ between models. Store these in the catalogue so a workflow cannot be pointed at a model that lacks what it needs.
Change one thing at a time
Move one workflow or team to the new model first, watch the logs and the evaluation results, then widen the change. The catalogue makes it just as easy to move back.
Decisions to make before you build
A managed gateway or your own
A managed product gives you routing, policy and logging quickly. Building your own gives full control, and the work of maintaining it. Either way, put it behind your own interface.
Where data is allowed to go
Decide which providers may receive which data before a model enters the catalogue. For some organizations this rules out particular providers or regions entirely.
What the same model means
Fix behaviour settings for each catalogue entry so a model behaves consistently across teams. Record the seed where results need to be reproduced.
Who owns the catalogue
Someone has to add, review and retire models. Treat the catalogue like any other shared configuration, with review and a history of changes.
Fallbacks and limits
Providers have outages and rate limits. Decide which requests may fall back to another model, which should wait and retry, and which should fail clearly. A fallback model needs the same evaluation as any other change.
What the logs may hold
Gateway logs and traces record prompts and responses, which can include client or personal data. Set who can see them, how long they are kept and what must be masked before anything is logged.
Cost by team and use
Because every call passes through one place, the gateway is where usage can be attributed to a team or workflow and limited. Agree the budgets and alerts before usage grows rather than after the first large bill.
The gateway as a dependency
Every model call now depends on the gateway, so it needs the same monitoring and recovery planning as any other critical service. A managed gateway also sees all of your model traffic, which belongs in the data decision above.
A checklist
- No workflow imports a provider SDK directly
- Every model is a catalogue entry with validated settings
- Provider keys live in the gateway or a secret store, not in applications
- Policies are defined once and attached by reference
- Every agent chain is traced, and logs follow agreed retention and masking rules
- A candidate model passes an evaluation set before it replaces the current one
- Fallbacks, rate limits and cost limits are decided per use
- The gateway product sits behind an interface you own
In practice
The case study below describes the platform this article draws on: the model catalogue, the two gateway options, guardrails attached per agent and tracing for agent chains.
One gateway to 45 models from six AI providers, with guardrails and tracing
A single model layer that lets teams choose and change models without rebuilding their workflows, with routing, policy and tracing handled in one place.
Read the case studyCommon questions
What is an LLM gateway?
A service that sits between your applications and AI model providers. It receives every model request, sends it to the chosen provider, applies policies such as guardrails and access rules, and records each call for debugging and oversight.
Do we need a gateway if we only use one provider?
Not always. But a provider-neutral interface and a model catalogue cost little to put in place early, and they turn adding a second provider into a configuration change rather than a project.
Does a gateway slow requests down?
It adds a step to every request. For most applications that step is small next to the time a model takes to respond, but measure it for uses where latency matters.
