5% of the GPU, All of the Bill: Running the Cloud Repatriation Decision on Numbers Instead of Ideology
Average GPU utilisation across 23,000 Kubernetes clusters is 5%, and 93% of enterprises are revisiting where AI runs. The answer is a utilisation number per workload, not a strategy slide.
The 5% Number That Ended the Cloud-First Default
Every infrastructure argument of the past decade eventually arrived at the same place: elasticity beats ownership, because you only pay for what you use. It was true, it built a trillion-dollar industry, and AI workloads have quietly broken the premise it rests on. You only pay for what you use is a statement about billing, not about consumption, and the two diverge badly the moment a workload needs an accelerator to be warm rather than busy.
Cast AI's 2026 State of Kubernetes Optimization Report put a number on it that is hard to argue with, because it is measured rather than surveyed: across 23,000 production Kubernetes clusters on AWS, Google Cloud and Azure, average GPU utilisation was 5%. Broken down by provider it was 2% on AKS, 5% on EKS and 6% on GKE. For context in the same dataset, CPU utilisation averaged 8% and memory 20%, which makes the accelerator, by a wide margin, the most expensive idle resource in the modern stack.
Attach a price to that idleness and the conversation changes tone. An H100 sitting unused on AWS p5 on-demand pricing runs at roughly $8,850 per GPU per month. A single eight-way node left running through a quarter of low activity is a six-figure line item that produced nothing, and Cast AI's own framing, that a twenty-engineer team with idle GPU habits can burn six figures a month without anyone noticing, is not hyperbole once you have seen how many development clusters are provisioned once and never scaled down. The same report noted something that has never happened before in this market: H200 capacity block prices on AWS went up in early 2026, the first GPU price increase since EC2 launched.
That last detail matters more than the utilisation figure, because it removes the assumption that has bailed out bad capacity decisions for fifteen years. Cloud unit prices always fell eventually. In a market shaped by the component shortage we covered in the hardware bill coming due, they no longer reliably do. If you were planning to grow into your cloud pricing, that plan now needs a fallback.
93%, 86%, 8%: Reading the Repatriation Statistics Honestly
The survey data arrived quickly and it is being quoted badly. Cloudian's Enterprise AI Infrastructure Survey, fielded in February 2026 across 203 enterprise IT decision-makers, found that 93% had already repatriated AI workloads from public cloud, were in the process of doing so, or were actively evaluating it. Seventy-nine percent had already moved something, 73% expected to shift further toward on-premises or hybrid over the following two years, and 86% expected AI budgets to rise in 2026, with 40% projecting increases of a quarter or more.
Two caveats belong in the same paragraph as that number, not three paragraphs later. Cloudian sells on-premises object storage, so the survey is not disinterested. And the 93% headline bundles together three very different populations: companies that have finished a migration, companies mid-project, and companies that once put repatriation on a meeting agenda. The third group is not evidence of anything except that the question is being asked.
The independent figures point the same direction with less drama. IDC found 86% of CIOs planning to repatriate at least some workloads in 2025, the highest reading it has recorded. Crucially, the same body of research puts full exits at 8-9%. Meanwhile the public cloud market kept growing at roughly 20% a year, which is only possible if new workloads are being created faster than old ones come home. The German analysis at Digital Chiefs is right to call the exodus framing a statistical illusion, and Gartner's projection that hybrid becomes the mainstream architecture for mission-critical work by the end of 2026 is the more accurate description of what is actually happening.
So the honest reading is narrower than the headline and more useful. There is no cloud exodus. What ended is the default. Between roughly 2015 and 2023, new workloads went to public cloud unless someone made a case against it, and that presumption did the deciding. In 2026 the presumption is gone, every workload with a meaningful compute footprint gets priced, and a minority of them come out the other way. That is a change in engineering process, not in ideology, and it is a change most organisations are not equipped to execute because nobody kept the skills to evaluate the alternative.
Key Takeaways
- Cast AI measured 5% average GPU utilisation across 23,000 production Kubernetes clusters, versus 8% CPU and 20% memory
- Cloudian's February 2026 survey: 93% repatriating, in progress or evaluating, from a vendor that sells on-premises storage
- IDC: 86% of CIOs repatriating something, but only 8-9% planning a full exit
- Cloud spend is still growing about 20% a year, so this is the end of a default, not an exodus
The Two Case Studies Everyone Quotes, and What They Actually Prove
Almost every repatriation argument in circulation leans on the same two examples, so it is worth reading them properly rather than as slogans.
37signals is the loud one. After announcing in October 2022 that Basecamp and Hey would leave AWS and Google Cloud, the company bought roughly $700,000 of Dell hardware and, as Data Center Dynamics reported, cut its annual cloud bill from about $3.2 million to $1.3 million, a saving of nearly $2 million in a single year. The Register followed the storage phase, in which the remaining roughly $1.5 million a year S3 bill went with the AWS account itself, taking projected savings past $10 million over five years.
GEICO is the more instructive one, because it is an enterprise rather than a software company with an unusually opinionated founder. The insurer began moving 600-plus applications to public cloud in 2013 and, by 2021, was spending north of $300 million a year across providers. Rebecca Weekly, its VP of platform and infrastructure engineering, has been unusually candid in public about the outcome, telling The New Stack that after a decade the bills had gone up roughly 2.5x and reliability had got worse rather than better. The response, covered by The Stack, was a large OpenStack private cloud with Kubernetes on bare metal underneath it.
What the two have in common is more diagnostic than the savings figures. Both run predictable, steady-state load rather than spiky consumer traffic. Both had, or deliberately built, a real platform engineering function with people who can operate hardware. Both had a decade of cost data to reason from, and both moved specific workload classes rather than performing a philosophical exit. Neither is a startup discovering product-market fit, and neither would advise one to buy servers.
There is also a survivorship problem worth naming, because it distorts every conversation on this topic. Nobody publishes the post-mortem of a repatriation that ran over, sat at 30% utilisation on hardware bought for a workload that got cancelled, and quietly went back to cloud eighteen months later. Those exist. They are just not conference talks.
The Break-Even Is a Utilisation Number, Not a Belief
Strip away the positioning and the decision reduces to arithmetic with three inputs: what the hardware costs to own and run, what the equivalent cloud capacity costs at the price you would actually pay, and what fraction of the year the workload genuinely needs the accelerator.
Work an example with the numbers above. Eight H100s on p5 on-demand at roughly $8,850 per GPU per month is about $70,800 a month, call it $850,000 a year to keep one node running continuously. Buying comparable eight-way capacity in 2026 lands very roughly in the $250,000 to $400,000 range depending on configuration and the state of the memory market, amortised over four years. Add power: a node of that class draws around 10kW, which at European commercial rates of €0.15 to €0.25 per kWh and a realistic cooling overhead comes to somewhere between €18,000 and €30,000 a year. Add colocation, networking, spares and support contracts, and an owned node in continuous use costs on the order of €120,000 to €160,000 a year all-in against $850,000 on-demand.
That ratio looks decisive, and it is also the most misleading number in this entire article, because almost nobody pays on-demand list for steady capacity. Apply a one or three-year commitment and the cloud figure falls by 40-60%. Now the comparison is much closer, and the deciding variable becomes utilisation: at 20% duty cycle the committed cloud rate wins comfortably, because you stop paying when the job ends and the owned node depreciates whether it is busy or not. This is why the published analyses cluster where they do. Lenovo's 2026 generative AI TCO analysis and similar independent work put the crossover somewhere between roughly 50% and 83% sustained utilisation against discounted cloud pricing, with small open models under 30 billion parameters reaching break-even fastest because they displace an expensive API call rather than a cheap GPU-hour.
The practical consequence is that the argument you keep having in the architecture channel is unanswerable without instrumentation, and instrumentation is the cheap part. Ninety days of per-workload utilisation telemetry costs almost nothing and settles the question definitively for each workload independently. The metric to build the case on is cost per successfully served request or per million tokens, not cost per GPU-hour, for exactly the reason we set out in the AI invoice nobody budgeted for: a GPU-hour dashboard will happily report a node that served nothing as fully productive.
Key Takeaways
- Break-even sits between roughly 50% and 83% sustained utilisation against discounted cloud pricing
- Benchmarking against on-demand list rather than committed rates overstates the on-premises case by 2-3x
- Small open models under 30B parameters cross over fastest, because they displace API pricing rather than GPU-hours
- Ninety days of per-workload utilisation telemetry settles the argument more cheaply than any consulting engagement
The Line Items That Never Make It Into the Spreadsheet
Hardware cost is the part of this decision that finance understands and engineering underestimates. The costs that actually sink repatriation projects are operational and they are all denominated in people.
Somebody has to own capacity planning, which in a cloud-native organisation has not existed as a discipline for years. Somebody has to run the procurement cycle, and GPU server lead times in 2026 still run weeks for standard configurations and considerably longer for volume orders, which means capacity decisions now need a two-quarter horizon rather than an API call. Somebody has to carry the firmware, driver and kernel compatibility matrix, which for accelerated computing remains genuinely unpleasant and breaks in ways that look like model bugs. Somebody has to design the failure domains, because a node with a dead NVLink is not a rescheduling event the way a failed cloud instance is. And somebody carries a pager for all of it, which means the real minimum viable team is three or four engineers, not one heroic infrastructure lead.
There is a hardware refresh mismatch underneath this that rarely gets modelled. A four-year depreciation schedule assumes the workload still wants that accelerator in year four. Model architectures and serving stacks have been turning over roughly annually, memory bandwidth requirements keep moving, and a card bought for one generation of inference workload can become poorly matched long before it is written off. That is a real risk, it is not a reason to avoid ownership, and it is a reason to prefer shorter horizons and stable, well-understood workloads for the first move.
Then there is the constraint that decides this for most European organisations regardless of the arithmetic: they cannot staff it. The Linux Foundation's State of Tech Talent Europe 2026 report puts Kubernetes and containers at the top of the reported capability gap list at 74%, and the shortage is sharpest exactly where repatriation needs it, in the second-order skills of platform engineering, reliability and infrastructure FinOps at senior level. Those people are scarce in Munich, Amsterdam and Stockholm, expensive when found, and typically three to six months from starting.
This is the point where a nearshore model stops being a cost story and becomes a feasibility one. At Stepto we staff dedicated platform and infrastructure engineers out of Serbia who work inside a client's own repositories, runbooks and on-call rotation, in a timezone that overlaps the entire European working day, which is what an infrastructure function actually requires and what an offshore rotation twelve hours away cannot provide. A hybrid estate needs the capability we described in platform engineering eating DevOps to exist permanently, not as a migration project, and building that capability is usually the binding constraint rather than the capital expenditure. We set out what that team costs, honestly, in our 2026 dedicated team cost breakdown.
Sovereignty and Latency Are Better Reasons Than Cost
Cost is the reason most repatriation business cases lead with, and it is the weakest of the three, because it is the only one that can be reversed by somebody else's pricing decision. A provider discount, a new instance family or a committed-use renegotiation can invalidate a spreadsheet that took a quarter to build. The other two reasons cannot be undone by a vendor.
The first is latency, and it is physics. Cloudian's survey found that 75% of respondents identified current or planned workloads that need or would materially benefit from on-premises infrastructure purely for latency: real-time video analytics, manufacturing quality control on a production line, low-latency transaction scoring. When an inference call has to complete inside the interval between two items passing a camera, the round trip to a region three hundred kilometres away is not a tuning problem. No amount of caching fixes the speed of light, and these workloads were never really candidates for centralised cloud inference in the first place.
The second is jurisdiction, and it is law. European organisations are now working through a stack of overlapping requirements about where data sits and who can be compelled to reach it, which we covered in the engineering requirements behind EU cloud sovereignty and, for regulated financial entities, in DORA's third-party risk regime. For a subset of workloads the placement question has already been answered by a regulator, and the only remaining engineering task is to build the thing properly rather than to litigate whether it is cheaper.
It is worth saying plainly that these two categories rarely describe the majority of an estate. They usually describe a handful of workloads, and those workloads are the correct place to start, because they are the ones where the decision is durable. Repatriating a stateless web tier to save 15% is a project you may well reverse. Repatriating a latency-bound inference service on a factory floor is a decision that will still be right in five years.
Key Takeaways
- Cost-driven repatriation can be reversed by a vendor's pricing move; latency and jurisdiction cannot
- 75% of surveyed enterprises identified workloads needing on-premises placement for latency alone
- Regulated and real-time workloads are the right first candidates because the decision is durable
Build So the Placement Decision Stays Reversible
The strategic goal is not to be on-premises or in cloud. It is to make placement a property of a workload that can be changed in a quarter rather than an architectural commitment that takes two years to unwind. Almost every organisation that got badly stuck in either direction got stuck for the same reason: the deployment target leaked into the application.
Kubernetes is the reason this is even tractable now, and it is genuinely the portability layer that makes the comparison practical across bare metal, private cloud and public cloud. It is also oversold. Kubernetes ports your workload scheduling; it does not port your managed control plane, your IAM model, your proprietary queueing and streaming services, your managed database's failover behaviour, or the GPU scheduling and topology awareness that took your team six months to tune. Those are where migration cost accumulates, and every one of them is a choice you make when you first build the service.
A workable discipline is narrow and boring. Use an S3-compatible object storage API so the storage tier is a configuration rather than a rewrite. Prefer portable data engines like PostgreSQL for anything whose placement might change, and reserve proprietary managed services for workloads you have consciously decided will never move. Describe infrastructure as code that expresses intent rather than provider primitives. Put model serving behind an OpenAI-compatible gateway so that swapping a hosted API for a self-hosted open-weight model is a routing change, which is the pattern we argued for in model portability as an exit strategy and the counterweight to the dependency risk in the AI infrastructure trap.
Watch data gravity above everything else, because it is the lock that actually holds. Compute is portable in a way that thirty petabytes are not, and egress pricing has historically been the mechanism that converted a cost argument into a captivity argument. The regulatory position here has improved for European organisations: under the EU Data Act, applicable since September 2025, cloud switching charges are being phased out and providers may not impose them from 12 January 2027. That materially changes the cost of changing your mind, and it is the single strongest argument for making the placement decision explicitly reversible in your architecture now rather than treating today's answer as permanent.
Key Takeaways
- Kubernetes ports scheduling, not IAM, managed data services or tuned GPU topology
- S3-compatible storage, portable data engines and an OpenAI-compatible serving gateway keep placement a config change
- Data gravity and egress are the real lock-in; the EU Data Act removes switching charges from 12 January 2027
How to Actually Run the Decision
A sequence that survives contact with a real estate looks roughly like this, and it deliberately starts with measurement rather than a business case.
Instrument first. Collect ninety days of per-workload utilisation, not per-cluster averages, because cluster averages hide the two services that consume everything and the forty that consume nothing. Then classify each workload along the axes that actually matter: steady-state or bursty, latency-bound or tolerant, regulated or not, data-heavy or not. Most estates resolve into a small set of steady, unglamorous, high-volume services that are genuine candidates, and a long tail that should obviously stay where it is.
Price honestly. Benchmark against the committed cloud rate you would negotiate, not the on-demand list price, and include four line items that repatriation business cases routinely omit: platform headcount, hardware refresh risk, procurement lead time, and the cost of running two environments in parallel during the transition. If the case only works at on-demand pricing or with zero incremental headcount, it does not work.
Then pilot exactly one workload. Batch inference, an internal service, or a stable high-volume API, with a defined rollback and a hard measurement window. Publish the real number afterwards, including the parts that went badly, because the organisational value of an honest first pilot is far larger than the savings on one service. Re-run the whole evaluation annually, since both hardware pricing and cloud pricing are moving faster now than at any point in the last decade.
This is the kind of work a dedicated team is well suited to and a project-based engagement is not, because it is continuous rather than finite. A Stepto team typically starts by instrumenting and modelling the existing estate before anyone commits to a direction, then carries whatever the numbers point to, hybrid orchestration, an on-premises inference tier, or a straightforward decision to stay in cloud and spend the effort on utilisation instead. That last outcome is a perfectly good result and, on the evidence of a 5% average GPU utilisation figure, it is the highest-return option available to a large share of organisations currently drafting a repatriation plan.
The Bottom Line
The repatriation story of 2026 is being told as a reversal and it is really a normalisation. Cloud-first was a useful heuristic during a decade when compute got cheaper every year and elasticity was worth paying for, and it stopped being a safe default the moment the dominant new workload class needed accelerators kept warm rather than instances scaled to demand. A measured 5% average GPU utilisation across 23,000 clusters, an idle H100 costing $8,850 a month, and the first GPU price increase in EC2's history are not an argument for buying servers. They are an argument for pricing each workload individually, which is exactly what almost nobody was doing. Do that and the honest answer for most of an estate is still public cloud with a serious utilisation programme attached, while a specific minority of workloads, the steady-state, latency-bound and jurisdictionally constrained ones, come out the other way and stay there. The organisations that will handle this well are not the ones with the strongest opinion. They are the ones with ninety days of real telemetry, an architecture where placement is a configuration rather than a rewrite, and a platform team that exists permanently instead of being assembled for a migration and disbanded afterwards. That team is usually the hard part to assemble, and it is the part worth solving first.
Building a team in Eastern Europe?
StepTo helps European and US companies build senior-led nearshore engineering teams in Serbia. Let's talk about what your next engagement could look like.
Start a conversationWritten by
Igor GazivodaCo-founder & CEO · StepTo
Igor has 15+ years in software engineering and business development. Former CTO at a Series A fintech startup, he specializes in scaling engineering teams, nearshore strategy, and AI-driven product development. He holds a Master's in Computer Science from the University of Belgrade and has published on distributed systems architecture.
LinkedIn →