The same open-weight model can run on your hardware, a managed provider or a private cloud tenancy. What each option costs, and what it gives back.

Most of the debate about open-weight models is about which one to choose. That is the easier half of the decision. The same checkpoint of GLM, Qwen or Mistral behaves identically whether it runs on a rack in the institution's own data centre, on a managed inference provider's shared GPUs, or inside a dedicated tenancy in a hyperscaler's cloud. What differs — sharply — is the cost curve, the data path, and who gets paged when it breaks.
That is why the hosting choice deserves as much attention as the model choice, and usually gets less. An institution can select an open-weight model precisely for sovereignty and then run it somewhere that quietly gives the sovereignty back.
Self-hosted. The weights run on hardware the institution owns or leases, inside its own network perimeter. No data leaves; no third party sees a prompt. The institution buys the GPUs, operates the inference stack, and owns the model lifecycle.
Managed inference provider. A specialist — Together AI, Fireworks, DeepInfra, Baseten and a growing field — runs the model on its own infrastructure and bills per token or per hour. The category has become serious money: Fireworks alone raised $1.5 billion this year. It offers most of the convenience of a frontier API while keeping the choice of model open.
Private cloud tenancy. The middle path: dedicated endpoints inside a hyperscaler or provider environment, often in a chosen region, with tenant isolation. Microsoft now offers open-weight models from DeepSeek, Qwen, Z.ai and others inside its Foundry catalogue, under Azure's governance and billing. Together AI offers dedicated EU-region endpoints on its higher tiers.
The honest case for a managed provider is strong, and for many workloads it is the right answer.
It buys the absence of an operating burden. There is no hardware to size, no inference server to tune, no capacity to plan for a quarter-end spike. For an institution without a machine-learning operations team, that absence is worth a great deal — running models is an operating commitment, and learning it mid-project is the expensive way. It buys speed to first deployment, measured in days rather than quarters. And it buys the ability to switch models as the rankings reshuffle without re-provisioning anything.
The catch is structural. The prompts, documents and outputs now travel to a third party again — which undoes part of the reason an institution chose open weights over Claude or ChatGPT in the first place. Region matters more than the marketing suggests: some providers run their serverless tier from US regions only, and EU-resident dedicated endpoints sit behind enterprise plans. Read where the inference physically happens before reading the price list.
Owning the stack inverts the trade. Data never leaves; the provider question disappears from the DORA third-party register; the model behaves identically in month one and month thirty because nobody else can update it. For a workload that a model risk function has validated, that permanence is worth more than it sounds.
The costs are equally real. Hardware is a capital commitment sized for peak load that sits partly idle off-peak. The skills — inference optimisation, monitoring, patching — are scarce and expensive. And the institution now carries the lifecycle: upgrades, capacity, failure. The hidden costs run in both directions, and self-hosting has its own list.
The lever that changes this arithmetic most is model size. A frontier-scale open model needs a cluster; a well-scoped small model doing one job runs on a single mid-range GPU. For bounded, high-volume work — the kind that makes up most production AI in a financial institution — self-hosting stops being an infrastructure project and becomes a server.
Institutions that work this through honestly tend to settle it on four questions:
Most institutions land in a hybrid: managed inference for experimentation and low-sensitivity work, self-hosted or dedicated capacity for the regulated production workloads where sovereignty and cost shape dominate.
If your institution has chosen its model and is now weighing where to run it, that is a conversation we have often.
Open weights give an institution the option to own its AI. The hosting decision determines whether it uses that option. A model downloaded for sovereignty and then served through a third party's shared GPUs is a frontier API with extra steps; a model sized to the job and run inside the perimeter is infrastructure the institution controls.
Neither is wrong. What goes wrong is choosing the model deliberately and the hosting by default.
Related reading:
Sources: provider documentation (Together AI, Fireworks AI, Microsoft Foundry), September 2026; funding as reported at the time of writing.
Stop renting generic models. Start building specialized AI that runs on your infrastructure, knows your business, and stays under your control.