Open-weight models vs Claude and ChatGPT: the real decision criteria

Where frontier APIs like Claude and ChatGPT still win, where open-weight models close the gap, and the five criteria that actually decide the choice.

Blog Collection Athour img
Michael Forystek
Co-founder, Growth & Partnerships
shape

Somewhere in every European financial institution this year, the same meeting is happening. One side of the table wants to keep building on the frontier APIs — Claude, ChatGPT, Gemini — because they are the best models available and the integration already works. The other side wants model weights inside the institution's own perimeter, for reasons that start with data and end with invoices. Both sides usually argue capability, which is the one dimension where the answer changes monthly.

The more durable way to frame the choice: this is not a comparison between models. It is a comparison between operating models — renting intelligence as a service versus owning it as infrastructure. Capability is one column in that table, and no longer the widest one.

The shrinking headline gap

The capability argument used to settle the meeting on its own. It no longer does. On Artificial Analysis's intelligence index, the best open-weight models now trail the proprietary frontier by six points, down from thirteen a year ago; on LMArena the Elo gap has compressed from roughly 150 points to about 30. The current open-weight leaders — DeepSeek's V4 line, the GLM-5 series, Qwen 3.5 — sit closer to the frontier than last year's frontier sat to itself.

The frontier keeps moving too. OpenAI shipped GPT-5.6 this month; Anthropic's Claude Opus 4.8 currently tops the same intelligence index. The gap is real, and it is now measured in months rather than years — which changes what an institution is paying for when it pays for the frontier.

What the API still buys

The honest case for Claude and ChatGPT is strong, and pretending otherwise is ideology rather than analysis.

Peak capability on open-ended work remains with the frontier. Tasks that are genuinely unpredictable — novel analysis, complex multi-step reasoning across unfamiliar domains, exploratory work whose shape changes weekly — reward the best available model, and the best available model is proprietary. The frontier also arrives first: new capabilities, longer context, better tool use land in the APIs quarters before equivalent open releases.

The API also buys the absence of a job. There is no hardware to size, no inference stack to operate, no model lifecycle to manage. For an institution without an MLOps capability, that absence is worth a great deal — running models is an operating commitment, and acquiring the skills mid-project is the expensive way to learn it. Elasticity matters too: a workload that spikes tenfold on quarter-end days is the API's natural territory.

If the workload is exploratory, modest in volume, and touches no sensitive data, the API is simply the right answer.

What the weights buy

Ownership pays in different currency. The first is sovereignty: with weights running inside the institution's perimeter, customer data, policy documents and case files never transit a third party. That simplifies the DORA third-party and concentration-risk conversation, shortens data-protection assessments, and changes the tone of the regulator's questions about where inference happens.

The second is cost shape. Per-token pricing is friction-free at pilot volume and compounds at production volume; owned infrastructure inverts that curve — the crossover arrives earlier than most projections assume for high-volume, always-on workloads like document processing or alert triage.

The third is permanence. API models get deprecated, re-tuned and re-priced on the provider's schedule; a validated model checkpoint on the institution's own hardware behaves identically in month one and month thirty. For systems that must be revalidated whenever the model changes — anything a model risk function has signed off — that stability is worth more than benchmark points.

The fourth is adaptation. Owned weights can be fine-tuned or RAG-refined on the institution's own documents, and a well-adapted smaller model routinely beats a general frontier model on the bounded, repetitive, domain-heavy tasks that make up most production work in a financial institution.

The columns that decide

Put capability in its place as one criterion among five, and the meeting gets shorter:

  • Task shape. Open-ended and unpredictable favours the frontier API; bounded, repeatable and document-heavy favours an adapted open model.
  • Volume economics. Estimate tokens per month honestly, then find the crossover. Pilots flatter the API; production volumes flatter ownership.
  • Data sensitivity. The more the workload touches customer data or regulated records, the heavier the sovereignty column weighs.
  • Internal capability. An institution with no ML operations capacity should price building it — or price a partner — before choosing ownership; an institution that already runs its own infrastructure loses the API's main advantage.
  • Time horizon. A quarter-long experiment belongs on an API. A capability the institution expects to run for years, under model-risk governance, argues for weights it controls.

Most institutions that work through the table honestly land in a hybrid: frontier APIs for exploratory and low-volume work, owned open-weight models for the regulated, high-volume production workloads where sovereignty and cost shape dominate.

Where this lands

The interesting shift of 2026 is that the choice has stopped being about whether open-weight models are good enough. For a widening class of production tasks, they are — which moves the decision from the benchmark table to the operating model, where it should have lived all along. Institutions that keep asking "which model is best?" will re-litigate the meeting every time a leaderboard updates. Institutions that ask "which column decides for this workload?" answer once.

The frontier is a subscription. The weights are an asset. Which one your institution needs depends on the workload in front of it — and working through that table for a specific use case is a shorter conversation than the meeting it replaces.

Related reading:

Sources: Artificial Analysis Intelligence Index (June–July 2026), LMArena, provider release announcements (July 2026). All figures as reported at the time of writing.

Ready to Own Your AI?

Stop renting generic models. Start building specialized AI that runs on your infrastructure, knows your business, and stays under your control.