Running open-weight models: self-hosted, managed provider or private cloud?

The same open-weight model can run on your hardware, a managed provider or a private cloud tenancy. What each option costs, and what it gives back.

Blog Collection Athour img
Michael Forystek
Co-founder, Growth & Partnerships
shape

Most of the debate about open-weight models is about which one to choose. That is the easier half of the decision. The same checkpoint of GLM, Qwen or Mistral behaves identically whether it runs on a rack in the institution's own data centre, on a managed inference provider's shared GPUs, or inside a dedicated tenancy in a hyperscaler's cloud. What differs — sharply — is the cost curve, the data path, and who gets paged when it breaks.

That is why the hosting choice deserves as much attention as the model choice, and usually gets less. An institution can select an open-weight model precisely for sovereignty and then run it somewhere that quietly gives the sovereignty back.

Three ways to run the same weights

Self-hosted. The weights run on hardware the institution owns or leases, inside its own network perimeter. No data leaves; no third party sees a prompt. The institution buys the GPUs, operates the inference stack, and owns the model lifecycle.

Managed inference provider. A specialist — Together AI, Fireworks, DeepInfra, Baseten and a growing field — runs the model on its own infrastructure and bills per token or per hour. The category has become serious money: Fireworks alone raised $1.5 billion this year. It offers most of the convenience of a frontier API while keeping the choice of model open.

Private cloud tenancy. The middle path: dedicated endpoints inside a hyperscaler or provider environment, often in a chosen region, with tenant isolation. Microsoft now offers open-weight models from DeepSeek, Qwen, Z.ai and others inside its Foundry catalogue, under Azure's governance and billing. Together AI offers dedicated EU-region endpoints on its higher tiers.

What the managed route buys

The honest case for a managed provider is strong, and for many workloads it is the right answer.

It buys the absence of an operating burden. There is no hardware to size, no inference server to tune, no capacity to plan for a quarter-end spike. For an institution without a machine-learning operations team, that absence is worth a great deal — running models is an operating commitment, and learning it mid-project is the expensive way. It buys speed to first deployment, measured in days rather than quarters. And it buys the ability to switch models as the rankings reshuffle without re-provisioning anything.

The catch is structural. The prompts, documents and outputs now travel to a third party again — which undoes part of the reason an institution chose open weights over Claude or ChatGPT in the first place. Region matters more than the marketing suggests: some providers run their serverless tier from US regions only, and EU-resident dedicated endpoints sit behind enterprise plans. Read where the inference physically happens before reading the price list.

What self-hosting costs

Owning the stack inverts the trade. Data never leaves; the provider question disappears from the DORA third-party register; the model behaves identically in month one and month thirty because nobody else can update it. For a workload that a model risk function has validated, that permanence is worth more than it sounds.

The costs are equally real. Hardware is a capital commitment sized for peak load that sits partly idle off-peak. The skills — inference optimisation, monitoring, patching — are scarce and expensive. And the institution now carries the lifecycle: upgrades, capacity, failure. The hidden costs run in both directions, and self-hosting has its own list.

The lever that changes this arithmetic most is model size. A frontier-scale open model needs a cluster; a well-scoped small model doing one job runs on a single mid-range GPU. For bounded, high-volume work — the kind that makes up most production AI in a financial institution — self-hosting stops being an infrastructure project and becomes a server.

The decision in four columns

Institutions that work this through honestly tend to settle it on four questions:

  • Data sensitivity. If the workload touches customer data or regulated records, self-hosting or a regionally isolated tenancy; if it touches only public or synthetic material, a managed provider is hard to beat.
  • Volume. Per-token pricing flatters pilots and compounds at production volume. Estimate monthly tokens honestly and find the crossover.
  • Internal capability. No MLOps team means pricing the build — or a partner — before choosing ownership. An existing infrastructure team removes the managed route's main advantage.
  • Regulatory exposure. A managed provider is an ICT third party with a register entry, a concentration-risk question and an exit plan to document. Self-hosting removes that entry and adds an operational-resilience obligation instead.

Most institutions land in a hybrid: managed inference for experimentation and low-sensitivity work, self-hosted or dedicated capacity for the regulated production workloads where sovereignty and cost shape dominate.

If your institution has chosen its model and is now weighing where to run it, that is a conversation we have often.

Where this lands

Open weights give an institution the option to own its AI. The hosting decision determines whether it uses that option. A model downloaded for sovereignty and then served through a third party's shared GPUs is a frontier API with extra steps; a model sized to the job and run inside the perimeter is infrastructure the institution controls.

Neither is wrong. What goes wrong is choosing the model deliberately and the hosting by default.

Related reading:

Sources: provider documentation (Together AI, Fireworks AI, Microsoft Foundry), September 2026; funding as reported at the time of writing.

Ready to Own Your AI?

Stop renting generic models. Start building specialized AI that runs on your infrastructure, knows your business, and stays under your control.