The Artificial Intelligence (AI) industry has remained fixated on model rankings. A new release claims the top of some leaderboard almost every week. Until recently, most enterprises simply chose the strongest available model and consumed it through managed APIs from the frontier labs. That decision is no longer straightforward.
What increasingly determines success is not which model scores highest, but which model — and which deployment approach — is right for a particular workload. Cost, governance, data residency, IP protection and operational complexity now sit alongside raw capability as first-order considerations.
The equation has changed
Open-weight models are the main reason that the choice has expanded. Unlike closed models delivered as a remote service, open-weight models allow organisations to run the trained weights themselves, subject to licence terms. Sensitive data can remain inside approved environments. Models can be fine-tuned on proprietary knowledge without routinely sending that knowledge to an external provider. Enterprises gain greater portability, reduce dependence on any single vendor’s road map and pricing, and often see substantially lower per-token costs — though total cost of ownership still depends heavily on utilisation and scale.
Open weights are not free. Downloading a model is the easy part. Operating it reliably at enterprise scale requires GPU infrastructure, inference serving, monitoring, security, governance, upgrades and licensing. Greater control comes with greater responsibility. For large organisations with deep engineering capacity, this trade-off can be worthwhile. For most mid-sized and small enterprises, it is far more challenging.
The July 2026 Hugging Face security incident illustrated the point with unusual clarity. When an AI-driven intrusion hit the company’s infrastructure, incident responders first turned to frontier models behind commercial APIs to analyse thousands of attacker actions. The forensic work required feeding real exploit payloads, attack logs and command-and-control artifacts to the models, but the providers’ safety guardrails blocked the requests; the systems could not distinguish an authorised responder from an attacker. Hugging Face completed the analysis on a self-hosted open-weight model instead. Sensitive incident data stayed inside its environment. The lesson was not that closed models are inferior. It was that some workloads structurally require a model you control. Security forensics, malware analysis and any investigation that must examine genuine attacker tooling cannot tolerate third-party guardrails that refuse the query or the risk of data leaving the organisational perimeter. Therefore, enterprises must classify workloads by control requirements as rigorously as by performance needs, and ensure that a capable, vetted open-weight model is already running on infrastructure they govern before an incident occurs.
Very few organisations have only one AI workload. A bank analysing confidential customer data has different requirements from a marketing team generating campaign content. A manufacturer embedding AI in customer service has different priorities from a cybersecurity team examining malware. Expecting one model and one deployment strategy to fit every use case is increasingly unrealistic.
The emerging third option
This is why managed inference platforms for open-weight models deserve more attention than another model launch. These platforms host leading open-weight families on controlled infrastructure and expose them through managed endpoints. Enterprises gain many of the benefits of open weights — data residency, fine-tuning flexibility and often lower cost — without having to build and operate the underlying GPU clusters and inference stack themselves.
Sarvam’s recent launch of Sarvam Inference, an India-hosted managed service unveiled at its Epoch 2026 conference, is one concrete example of the category taking shape. The platform currently serves Sarvam’s own 105-billion-parameter model alongside leading open-weight families such as GLM 5.2 and Gemma 4, all running on domestic infrastructure. The significance is not any individual model. Enterprises could already download many of them. The challenge was making them work reliably in production — handling concurrency, latency, security and continuous updates at scale. Managed inference changes that equation. By offering production-grade endpoints under Indian data residency, it is likely to democratise access for companies that could never justify specialised AI operations teams while supporting the broader push for “token sovereignty.”
One caveat remains important. Managed open-weight platforms reintroduce vendor dependence — at the infrastructure layer rather than the model layer. Enterprises should evaluate portability guarantees, security posture, pricing trajectory and exit paths with the same rigour that they apply to any frontier API contract.
It is about deployment choice
The way forward is to match the workload, not the leaderboard. No single deployment model is right for every organisation or every use case. Customer-facing tasks that demand frontier reasoning often fit closed APIs. Regulated workloads with strict data-residency obligations frequently suit managed open-weight platforms hosted in-country. Security forensics, malware analysis and IP-critical fine-tuning usually belong on self-hosted deployments.
The discipline lies in making that call workload by workload rather than by corporate default. Companies that invest in understanding the strengths, limitations and economics of each approach — and that systematically match every workload to the option delivering the right balance of capability, control, cost and governance — will extract far more value from AI than those still chasing the latest leaderboard ranking. The competitive advantage will belong to those who treat deployment choice as a core architectural decision, not a procurement afterthought.
Chandrajit Banerjee, Director General of the Confederation of Indian Industry; Debjani Ghosh is Distinguished Fellow – NITI Aayog and former President, Nasscom
Published - August 18, 2026 12:08 am IST