When AI infrastructure becomes expensive, the first instinct is often to look at the GPU bill. But in many production environments, the more important question is what those GPUs are waiting for.
Consider Synthesia, an enterprise AI video communications platform that uses large-scale computing to develop increasingly realistic text-to-video models. As its AI workloads grew on AWS, the company encountered a problem that had little to do with the raw capability of its GPUs. Legacy storage and data-management processes were creating bottlenecks, resulting in low GPU utilization and slowing model training. Synthesia ultimately moved to a data platform designed to keep training data available to compute while reducing manual data movement and simplifying migration across AWS Regions.
So, AI performance is not determined by compute capacity alone. It depends on how efficiently data can reach that compute, how often it needs to move, where it is stored, and what it costs to keep the entire pipeline running. This is where data gravity becomes more than an architectural concept, It becomes an economic decision.
The GPU Is Only Productive When the Data Can Keep Up
That economic shift becomes particularly important as organizations move AI workloads from experimentation into production. Modern GPUs can process enormous volumes of data, but that capability creates a new infrastructure requirement: storage and networking have to deliver data at a rate that keeps the accelerators busy. HPE and NVIDIA have highlighted this shift directly, noting that data movement, locality and pipeline efficiency can become limiting factors even as GPU and networking performance continues to increase.
The implication for infrastructure leaders is important. Buying additional GPUs does not automatically increase useful AI capacity. If the storage layer cannot deliver data fast enough, expensive compute sit underutilized.
Technologies such as NVIDIA GPUDirect Storage illustrate why the path between storage and GPU matters. By reducing unnecessary CPU involvement in the data path, GDS can improve throughput and reduce bottlenecks; NVIDIA has reported more than 2x gains in some configurations and substantially higher improvements in specific storage workloads.
Once the data path becomes part of the performance equation, another question follows naturally: What happens when that data has to travel across infrastructure boundaries?
Data Movement Has a Price
The financial impact becomes more visible when large datasets begin moving between clouds, regions, or environments.
Cloud providers charge for many forms of data transfer, particularly data leaving a cloud environment. Microsoft Azure, for example, lists internet data transfer rates that can reach approximately $0.12 per GB for certain geographies and routing scenarios after the applicable free allowance. At that rate, moving 100 TB could represent roughly $12,000 in transfer charges alone. That calculation does not include storage, networking infrastructure, synchronization, processing, or the operational cost associated with managing additional copies.
Google Cloud similarly applies data-transfer charges that vary by destination and volume, with some outbound traffic priced around $0.12 per GiB at lower usage tiers. A single transfer may be manageable but the problem emerges when the same dataset moves repeatedly between regions, clouds, development environments, training clusters, analytics platforms, and disaster recovery locations which ultimately changes the infrastructure equation thus, the cost of AI is no longer simply the price of GPUs and storage, It includes the cost of getting the right data to the right compute at the right time.
Data Gravity Changes the Placement Decision
Data gravity describes the tendency of large, valuable datasets to attract applications, services, and compute toward them. In an AI environment, that pull becomes stronger because training and inference can involve enormous datasets and frequent data access.
This does not mean enterprises should automatically move compute on-premises. Cloud GPUs can still make strong economic sense for workloads that are temporary, unpredictable, or require rapid scaling.
The question becomes more specific:
Where should the workload run when the cost of accessing its data is included?
For an occasional training job, paying for cloud compute and transferring a dataset may be reasonable. For continuous inference against a large and frequently changing enterprise dataset, repeatedly moving information to distant compute may create unnecessary cost and latency.
The decision should therefore account for the workload’s actual behavior: dataset size, access frequency, rate of change, data sensitivity, latency requirements, workload duration, and how often information needs to cross cloud or regional boundaries.
Once those factors are considered together, one issue becomes particularly easy to overlook.
The Hidden Cost Is Often Duplication
Data movement does not always appear as a line item labelled “data movement.” It often appears indirectly through duplicated datasets. A copied dataset is not simply a storage problem; it creates a synchronization problem. The more copies exist, the harder it becomes to establish which version is current, which one is governed, and which one can be recovered reliably.
For enterprises operating AI at scale, reducing unnecessary copies can therefore be as valuable as reducing raw storage consumption.
Hybrid Cloud Becomes a Workload Placement Strategy
Hybrid infrastructure is often discussed as a compromise between on-premises and cloud. For AI, it can be more useful to think of hybrid architecture as a workload placement strategy.
Data that is sensitive, highly regulated, or accessed continuously may make economic and operational sense closer to controlled infrastructure. Bursty workloads can use public-cloud GPU capacity. Frequently accessed datasets can remain close to the compute that uses them most. Less frequently accessed information can move to lower-cost storage tiers.
The objective is not to choose one environment for everything. It is to reduce unnecessary distance between compute and data while preserving flexibility.
That approach is already visible in AI infrastructure design. Oracle’s work with MosaicML, for example, uses high-bandwidth networking and OCI storage to stream large datasets efficiently to GPU infrastructure without the same type of cross-cloud egress considerations (Oracle, 2025).
The strategic question is therefore shifting from “Cloud or on-premises?” to “Which environment is economically and operationally appropriate for this workload and its data?”
What Should Technology Leaders Actually Measure?
GPU utilization remains important, but it is only one part of the equation. A more complete view of AI infrastructure economics should include:
GPU utilization + data-access latency + storage throughput + transfer volume + dataset duplication + synchronization overhead + operational cost.
Consider two environments with identical GPU capacity. If one consistently keeps GPUs supplied with data while the other experiences storage bottlenecks, network delays, and repeated dataset transfers, their real cost per workload will be very different.
This is why infrastructure teams should begin measuring the total cost of the AI data path, rather than looking at compute pricing in isolation.
The New Infrastructure Equation
Taken together, these factors point to a broader change in how AI infrastructure should be designed. Compute, storage, and networking can no longer be evaluated as completely separate layers
The answer is also not that every organization should move GPUs closer to its data. The right decision depends on the workload. Some environments will benefit from keeping compute close to large, continuously accessed datasets. Others will benefit from moving data efficiently to elastic cloud compute. Still others may require a hybrid model that combines controlled data environments with cloud-based acceleration.
At Open Storage Solutions, we see storage as an increasingly active part of that infrastructure equation. As enterprises evaluate AI, hybrid environments, and growing data volumes, the objective is not simply to store more data, but to build an architecture that allows data to remain accessible, resilient, and available where workloads need it.
So, before investing in more compute, understand the cost and performance of the data path feeding it. As, the future of AI infrastructure is not simply about moving compute to data or data to compute. It is about minimizing the expensive distance between them.
Add your first comment to this post