Not every AI workload belongs in the same place. Large language models often fit logically in the cloud because they can serve as general-purpose engines that improve with scale and draw value from broad, up-to-date knowledge.
But the AI workloads moving into production are not just text and language, and we’re increasingly seeing enterprises adopt multimodal models for AI video, audio, and image generation and seeing massive advantages across compute usage, control of IP, and ability to customize the look and feel of creative output.
Co-founder and CEO of LTX.
For these types of creative production, the raw material they’re using to build is not the open web. Instead, it’s often footage, branded assets, or unreleased IP that already lives within the organization’s walls. In these instances, there’s a clear need for running the models closer to where that content already resides.
That’s because creative production is iterative by nature, and that volume of iteration and generation brings with it real cost pressures when drawing on the cloud.
As a founder, I’ve watched this transition play out repeatedly: companies adopt AI pilots, usage skyrockets, and suddenly finance teams are trying to understand which teams, workflows, or model calls are driving up the bill.
Cost predictability becomes an infrastructure question
Once AI tools become part of daily work, usage no longer behaves like an experiment. Every generation, agent action, video render, or workflow step carries a cost. The equation becomes much harder to forecast once adoption spreads across teams and AI agents.
For companies with high-volume creative workloads, running more of their inference locally, at the edge, or in private environments gives greater control over unit economics and makes AI spending easier to manage over time.
This is particularly important in creative production environments, like filmmaking and gaming, all the way to marketing campaign creation and internal training, where teams often generate dozens of variations of an asset, sequence, campaign concept, or interface.
In an environment where a single workflow can generate thousands of API calls per day, the difference between cloud and local inference can determine whether an AI strategy is sustainable or requires constant budget justification.
Data control will shape deployment choices
Long-term, data control has potential to be a primary driver for enterprises to move toward more flexible AI architectures. Businesses have become increasingly sensitive about where and how their information is stored, how long it stays there, who has access to it, and how it can be used.
Those questions become more serious when AI is mapping physical environments, working with unreleased creative assets, production files, or other material that was never meant to move freely outside controlled systems.
When it comes to AI video generation, which can involve multiple iterations on sensitive creative assets and IP, teams may prefer to run their models within their own environments. In these cases, local or private deployments are less about rejecting the cloud and more about giving companies a way to use AI without handing over access to sensitive information.
As AI becomes more embedded in business-critical work, these choices will involve more than IT architecture because they affect what a company can build, what risks it takes on, and how much control it keeps over the systems producing its work.
The future is optionality, not a single deployment model
The cloud has proven to be essential for many AI workloads, especially when companies need elastic compute, access to frontier models, or the ability to support highly variable demand.
A more realistic future is one in which enterprise AI becomes hybrid by necessity, with different workloads running in different environments based on the needs of the business rather than the convenience of a single deployment model.
Some workloads will run in the cloud because scale matters most, while others will run locally because latency, interactivity, and iteration matter more, and still others will run on-prem to prioritize privacy, compliance, customization, or ownership.
The organizations that prepare for this transition will be the ones that stop treating deployment as a binary choice and start asking which workloads require which level of control.
This pressure only intensifies when we consider where creative production is heading. The same models that teams use to generate video are now evolving into world models: systems that can predict and simulate the physical world, moment to moment, in real time.
Workloads like these will be defined by interactivity and latency, and a generation that waits on a round trip from the cloud and back won’t be able to cut it.
This article was produced as part of TechRadar Pro Perspectives, our channel to feature the best and brightest minds in the technology industry today.
The views expressed here are those of the author and are not necessarily those of TechRadarPro or Future plc. If you are interested in contributing find out more here: https://www.techradar.com/pro/perspectives-how-to-submit