For the last two years, most enterprises have bought AI the same way: by the token, through someone else’s API. It was the fastest way to start. It’s also the fastest way to build a cost structure you don’t control.
There’s a second option now, and it’s grown up fast: download the weights, run the model yourself, tune it for your business. Here’s why that matters, and where the real distinctions are.
Closed-weight, open-weight, open-source: the terms get blurred
Closed-weight is the model behind the vendor’s API. OpenAI’s GPT, Anthropic’s Claude, Google’s Gemini. You rent inference by the token. The vendor controls the weights, the pricing, and the deprecation schedule. When they retire a model, your integration retires with it.
Open-weight means you can download the actual weights and run them yourself. Llama, Mistral, Qwen, DeepSeek, Kimi, GLM. But “open” usually means the weights only. The training data stays secret, and the license may still restrict what you can do commercially.
Open-source is weights plus an OSI-approved license — MIT or Apache 2.0. MiMo-V2.6-Pro and DeepSeek-V4.1-Flash ship under plain MIT. Most models people call “open” are open-weight, not open-source. Check the license before you build on it. The distinction matters the moment lawyers get involved.
Benefits unique to open-source
1. License freedom. MIT or Apache 2.0 means you can modify, redistribute, and embed the model in commercial products without legal exposure. No bespoke license restrictions, no per-company terms, no acceptable-use policy that changes under you. Your legal team reviews a standard license they’ve seen a thousand times, and you ship.
2. Real auditability. With the weights, code, and ideally the data pipeline visible, your security team can actually inspect what’s inside the model instead of trusting a vendor’s safety report. In regulated industries, “trust us” is not a control. Inspectability is.
3. Forkability. The community can fork, improve, and maintain the model. You’re not dependent on one vendor’s roadmap. If the original author walks away or changes direction, the model doesn’t die with them. That’s a continuity guarantee no SLA can match.
Benefits of open-weight or open-source
1. You control the cost curve. Host on your own infrastructure and your team can push token volume freely — token maxxing without a meter running. No surprise usage bill at the end of the month, no sudden API price hike, no dependency on a model that gets deprecated out from under you. A Mozilla analysis found open models within benchmark noise of closed alternatives on standard workloads, at 40–70% lower cost. At enterprise token volumes, that gap is a budget line.
2. You protect your alpha. Whatever differentiates you — your data, your processes, your customer knowledge — stays inside your perimeter. Every prompt you send to a third-party API is data leaving the building. Self-hosting eliminates that exposure entirely. For companies whose edge is proprietary knowledge, this alone justifies the move.
3. You tune for your domain. Small open models fine-tuned with techniques like QLoRA can match or beat frontier closed models on specific domain tasks, at a fraction of the operating cost. A 70B model tuned on your data, your terminology, your workflows will outperform a giant generalist on your work. Generic models answer like everyone else’s. Yours answers like you.
4. Portability. Benchmark models head-to-head on your own data. Swap components without refactoring your stack. And negotiate from strength, because your exit path is real. Nothing disciplines a vendor like a customer who can leave.
5. Compliance by design. Data residency and audit requirements get simpler when the model lives where the data lives. Inference logs stay in your perimeter. You enforce your own change control, your own red-teaming, your own review gates — under frameworks like NIST’s AI RMF — instead of inheriting someone else’s.
The field right now: a snapshot, not a ranking
Model leadership changes quarterly. Treat the table below as a snapshot of October 2026 — representative models in each tier, not a buying guide for 2027. For the current picture, check the leaderboards that update continuously:
- Artificial Analysis — independent benchmark scores across quality, price, and speed, updated as new models release. Start here for price/performance on your workload type.
- LMArena — crowd-voted head-to-head model rankings. The closest thing to a public preference poll for which models people actually prefer.
- OpenRouter Rankings — which models are getting real API traffic and token volume in production right now. Shows what the market is running, not just what’s marketed.
How to use them together: find candidate models for your workload on Artificial Analysis, cross-check real-world preference on LMArena, then confirm production traction on OpenRouter — and always read the actual license on the model card before you commit.
| Model | Tier | License | Known for |
|---|---|---|---|
| OpenAI GPT-5.5 | Closed-weight | Proprietary | Frontier reasoning; leads on agentic benchmarks |
| Anthropic Claude Opus 4.7 | Closed-weight | Proprietary | Coding and complex agentic work |
| Google Gemini 3.8 | Closed-weight | Proprietary | Multimodal; very long context |
| Anthropic Claude Fable 5 | Closed-weight | Proprietary | Knowledge work; strong reasoning |
| xAI Grok 4 | Closed-weight | Proprietary | Reasoning with real-time information access |
| DeepSeek V4 | Open-weight | Open weights, custom license | Price/performance leader; huge production token volume |
| Kimi K3 (Moonshot AI) | Open-weight | Open weights, custom license | Near-frontier quality at a fraction of frontier price |
| Qwen (Alibaba) | Open-weight | Open weights, custom license | Strong multilingual; broad range of model sizes |
| Llama (Meta) | Open-weight | Open weights, custom license | Widest ecosystem and tooling support |
| GLM (Zhipu AI) | Open-weight | Open weights, custom license | Strong agentic and coding performance per task cost |
| MiMo-V2.6-Pro (Xiaomi) | Open-source | MIT | Frontier-class capability under a plain MIT license |
| DeepSeek V4.1 Flash | Open-source | MIT | Efficient high-volume model under MIT |
| OLMo (Ai2) | Open-source | Apache 2.0 | Training data and code fully open |
| Granite (IBM) | Open-source | Apache 2.0 | Enterprise-focused models under Apache 2.0 |
| Phi (Microsoft) | Open-source | MIT | Small efficient models for edge and device use |
Freshness disclosure: this table was accurate in October 2026 and will not stay that way. New models ship monthly, licenses change, and today’s leader is next quarter’s baseline. Before choosing a model, verify the current rankings above and read the license on the model’s own page — not a blog post, including this one.
The honest caveat
The infrastructure burden is real. GPUs, MLOps, patching, scaling, monitoring — that’s all yours now. Self-hosting trades a vendor bill for an operations commitment, and you should price that honestly before you commit.
And for the hardest agentic work — long-horizon reasoning, million-token context, complex tool use — closed models still hold a real lead, roughly a four-month capability gap per Epoch AI. If your use case lives at the frontier, rent the frontier.
Open weights aren’t a religion. They’re an architecture decision. Make it workload by workload.
The bottom line
But for most enterprise workloads, the math has flipped. The models are good enough, the tooling is mature enough, and the cost and control advantages are large enough that the default is shifting from rent to own.
The question isn’t whether you can afford to host your own models.
It’s whether you can afford not to.