Over the past few months, almost every project brings up the same conversation. The customer looks at the estimate for two sites, sees the total, and asks: what if we skip building the second site and rent one instead? Sometimes it goes the other way — they show up with a ready-made plan to rent everything, and we sit down to work out what that will cost by the end of the second operating cycle.
The trouble with this conversation is that it’s treated as one question, when it’s really four. Core virtualization, backup, analytics, and AI workloads are built so differently that a single answer doesn’t fit all of them. Renting, which looks reasonable for training models, becomes an expensive mistake for round-the-clock production. Owned hardware, mandatory for a bank’s core, turns out to be overkill for the second tier of backup copies.
Below are the four circuits and where the line falls in each. Half of this piece is about physics and arithmetic: it’s usually those, not the price list, that settle the argument. The other half is about what changed in 2026 — in component prices and in the regulatory framework.
Circuit one · Distance
Distance decides before the budget does
5 microseconds per kilometer — a number no budget or contract can get around.
Start with the kilometers, not the money. Site separation sets everything else, and it sets it hard.
Light in fiber travels at roughly 200,000 km/s, which works out to 5 microseconds per kilometer. The actual route is always 30–50% longer than the straight-line distance — terrain, existing ducts, detours. That gives a simple rule: roughly 1 ms of round-trip time for every 100 km of route. 200 km in a straight line comes to about 4.2 ms, 400 km to 6.9 ms, 600 km to 9.6 ms.
Now let’s line this up against platform requirements. For stretched clusters on leading virtualization platforms, the ceiling between data sites is 5 ms round-trip, i.e., 2.5 ms one way. That’s not an arbitrary limit: the application has to manage to write data to both sides in time. Metro clusters on external arrays sit in the same range. So 200 km is right at the edge, and 400 and 600 km fall out of a synchronous scheme entirely.
This is the conclusion worth spelling out to the customer before pricing the spec: at 400–600 km, there’s no such thing as a “full hot-standby cluster copy.” What you get instead is asynchronous replication with a recovery point measured in minutes and a failover that someone has to trigger. That’s a different architecture, a different runbook, and different money.
A separate note on the regulator. Regulation No. 3669 of August 18, 2025, on minimum information security requirements for commercial banks — in force since November 20 — requires the backup data center to sit at least 50 km from the primary one, and mandates that information assets, security controls, databases, and servers be located at the bank’s primary and backup centers: in the bank’s own building, its branches, the Central Bank’s cloud data center, or state-owned data centers.
50 km is under 1 ms — synchronous writes hold up fine. That gives an honest fork in the road, and we prefer to put it on the table openly: either 50–100 km with a zero recovery point, but both sites in the same seismic and power region, or 300–600 km with real protection against a regional event, at the cost of synchrony. Both options are compliant. The first is usually chosen by those who fear an outage; the second, by those who fear an event.
And the second half of the same point: the regulation’s list of permitted sites is closed. That means a bank’s rental scheme has to be checked against that list first, and only then against the cost calculator. This isn’t a ban on renting — it’s a requirement to know who you’re renting from.
Circuit one · The price of a duplicate
What the second copy really costs
The backup site doesn’t just pay in hardware: cores are licensed on both sides, and perpetual licenses aren’t sold anymore.
This is where the arithmetic starts, and it’s where customers regularly come up half short.
Take the platforms in common use today. A leading hyperconverged infrastructure vendor states its licensing policy outright: disaster recovery sites require separate licensing, and licenses for file and object services are needed on both the source and the target, sized to the capacity consumed on each side. In other words, the backup site pays almost as much as the primary one — and not only in hardware.
On the classic three-tier stack, the story isn’t any better, and it got noticeably worse after the platform changed owners in 2024. Licensing is now subscription-based and counted per physical core, with a minimum of 16 cores per processor — an 8- or 12-core processor is still billed as 16. Every core on every host counts. A backup site on the classic stack means another set of servers with their own cores, another array, another storage fabric, and another subscription on top.
Hidden in here is what we’d call the biggest news of the last two years for the whole “rent or own” question. Buying used to work like this: buy the hardware, buy a perpetual license, lock in the platform’s cost for its entire service life, and pay only for support after that. Today there are no perpetual licenses on new contracts. You buy the hardware, but you rent the platform either way. By consultants’ estimates, moving from a bare hypervisor to a full bundle pushed annual costs up 2–5x, and the per-core minimums added their own share on top.
Which leads to a conclusion few people like: buying hardware no longer protects you from rising software prices. The argument worth having now isn’t buy versus rent — it’s which part of the stack you’re willing to keep on a subscription, and for how long you’ve locked in the price.
But over the past year, a second factor has been added that flipped the estimate’s structure right back.
Circuit one · Component prices
Hardware is the biggest line item again
While everyone was busy counting the license perimeter, the estimate moved over to hardware — and memory is to blame.
As recently as 2024, a typical HCI project split roughly evenly between software and hardware. Today, in projects of the same kind, we see a completely different ratio — about 20% on licenses and about 80% on hardware. Especially wherever the design carries a lot of RAM and flash in any form: NVMe, SSD, cache tiers, backup shelves.
The cause isn’t licensing — it’s memory. In April 2026, Gartner forecast average DRAM prices rising roughly 125% over 2026, NAND roughly 234%, with no meaningful relief before the end of 2027. Quarterly figures from TrendForce make it even clearer: in Q1 2026, contract prices for commodity DRAM rose 90–95% quarter over quarter, in Q2 another 58–63%, and in Q3 growth slowed to 13–18% — but off an already inflated base.
It matters to understand the nature of this increase, or you’ll end up waiting for a correction that isn’t coming. This isn’t a COVID-style supply shortage that clears up in a quarter. It’s a structural reallocation of fab capacity: wafers are shifting toward memory for accelerators and enterprise NVMe drives for data centers, while hyperscalers buy up available volume on long-term contracts. IDC estimates DRAM supply growth for 2026 at just 16%, NAND at 17%, against a historical norm of 20–30%. New fabs won’t reach full volume before the end of 2027.
Three things follow for our purposes, and none of them boil down to “it costs more to buy now.”
First: the point of comparison has shifted, but in both directions at once. Buying got more expensive — that’s obvious. But the provider buys the same hardware from the same suppliers, so rental rates will follow, with a lag equal to the length of their current contracts. A comparison that uses last year’s rental rates against today’s purchase prices lies in favor of renting.
Over two quarters, memory got 3.1x more expensive, drives 2.7x. The slowdown is measured from that base — not from the one your last year’s budget was built on.
Second, and this matters more than price: lead times for large memory orders have stretched past 40 weeks. The risk now isn’t that a configuration gets more expensive — it’s that the modules and drives you need may simply not be available in time, at any price. That’s a separate risk line, and it needs discussing before signing, not after.
Third: rising prices hit the backup site harder than the primary. The primary is usually already bought; the backup still has to be purchased — at the new prices. That’s the most practical argument for a partial standby, which we get to next.
And the takeaway for this section: since hardware’s share of the estimate has grown to 80%, the ceiling on savings from the license perimeter is one-fifth of the budget. The other four-fifths are determined by the memory and flash configuration.
Circuit one · Standby service level
Standby doesn’t have to be full
The service level on standby is a decision you make and write down — not something that just happens on its own.
This is the single biggest saving in this whole conversation, bigger than any rental decision, and for some reason it’s the one discussed least.
If you describe the degradation up front — on the backup site we bring up Category 1 systems, and leave Category 2 and 3 down until the primary is restored — the standby runs comfortably at 50–60% of capacity. That’s not cutting corners; it’s a deliberate service level, written into the runbook and verified in drills. Cutting corners is when a 100% standby is drawn up on paper, and the drill reveals that half the systems on it won’t even start.
This is also the place to close off a temptation that hits everyone who’s seen component prices: putting hardware that’s finished its first cycle on the primary site onto standby. We plan hardware for 3 years of primary operation under warranty plus 3 years of a second cycle on post-warranty support, and the arithmetic looks tempting.
For production workloads, we’re against it, for two reasons. The first is organizational: a site built from old hardware stops being taken seriously, and within six months people quietly stop caring for it — firmware doesn’t get updated, drills don’t happen, disk degradation goes unwatched. The second is technical: second-cycle hardware has a noticeably higher chance of failing at exactly the moment load hits it — that is, at or just before the outage on the primary site. A standby that breaks when it’s asked to take the load is worse than no standby at all — people were counting on it.
On top of that, today’s norm in banks is matching hardware on both sites, a requirement from risk departments that’s usually written down explicitly. Arguing against it with technical reasoning is pointless and unnecessary.
For test environments, development, and workloads whose downtime the business can live with, second-cycle hardware is a perfectly fine choice. That’s exactly where it should be redeployed, freeing up budget for the production standby.
The third move is to take out of the standby anything that doesn’t physically need to sit there. The next two sections cover that.
Circuit two · Link arithmetic
Link arithmetic
What travels over the inter-site link isn’t the full volume — it’s the unique blocks. There are exactly two bottlenecks, and both deserve their own line item.
There’s a persistent myth about backup that we repeated ourselves for a while: that inter-site replication is bottlenecked by the link, because “you can’t move terabytes a day.” Let’s do the math, because the numbers turn out to be surprising in both directions.
First, the upper bound. At a realistic 85% useful utilization, a 100 Gbit/s link moves about 300 TB in 8 hours and about 900 TB a day. A 10 Gbit/s link moves 31 TB in 8 hours, 92 TB a day. This immediately shows that the claim “move 1,000 TB overnight on a hundred-gig link” is wrong: 1,000 TB at 100 Gbit/s takes 26 hours — it won’t fit in an overnight window.
But that’s not the main point. Nobody pushes the full volume across the inter-site link. What moves is the changes, and after deduplication, only the unique blocks. Take a setup with 1,500 TB of protected data, a daily change rate of 2–5%, and a wire-level compaction factor of 3 to 5x. That comes out to 9–38 TB a day — 2–10 hours at 10 Gbit/s, under an hour at 100 Gbit/s.
So for day-to-day operation, the link isn’t the bottleneck — provided deduplication actually works, and the receiving end at the second site can ingest unique blocks rather than a full stream.
But this picture comes with an important caveat, and you need to check it against your own policy, not the market average. The delta stays small as long as incremental jobs are running. The moment a scheduled full copy is required — and many policies mandate one weekly or monthly — the daily volume jumps by an order of magnitude and runs straight into the hours from the table above. So the number to plan around isn’t the average day, but the worst one: the night after a full backup, with every other job running at the same time. If the window doesn’t close on that day, it doesn’t close at all.
The link becomes a bottleneck at exactly two points, and both deserve their own line in the project. The first is the initial seed, when the copy has to be created from scratch. The second is mass restore, when data has to be pulled back. Everyone remembers the first; people remember the second only once the outage happens.
Circuit two · Storage layout
A second site without a second full set
A full three-tier complex on the primary site; media servers and rented storage on the backup one.
The layout we arrive at when pricing this kind of second site looks like this.
On the primary site, we build a full three-tier complex: a fast NVMe tier for quick restores, a disk tier for the main retention depth, and a tape tier for long-term storage and immutable copies. All owned, all under our control, restore from the first tier measured in minutes.
On the second site, we don’t build a full duplicate of that setup. We deploy media servers and rented object storage, with a small NVMe array nearby as a landing zone. The argument here isn’t about the network, and that matters: data going into object storage also travels over the network — the same traffic. The argument is capital: you don’t need a second set of disk shelves and a second tape library, and that’s the dullest and most expensive part of a second site. Given what’s happening to drive prices, the savings here have grown noticeably over the past year.
Now, honestly, the downsides — because this scheme has them.
The return path costs money and time. Upload traffic to the cloud is usually free; download isn’t, and on a mass restore, egress traffic becomes its own line item. Cold storage tiers carry a minimum billable duration: delete an object early, and you still pay for the full term. For short retention, that breaks the economics completely, so retention depth needs to be split by tier: 7–14 hot days in a fast tier or locally, everything else in an archive tier.
The NVMe landing zone isn’t a luxury — it’s a mandatory part of the scheme. Without it, restore time is set by the return link to the cloud, and we’ve already worked out how many hours that means. With it, you pull down what you need ahead of time or in batches, and restore from fast local storage instead.
And for banks — check against Regulation No. 3669’s list of permitted sites. A backup of the core system contains information covered by banking secrecy, and the question of who you’re renting the object layer from is regulatory here, not commercial.
Circuit three · Data platform
Analytics has four layers, not one
The requirement to “replicate the whole lake” is almost always overkill. And the most critical layer turns out to be the smallest.
With data platforms, the situation runs the other way: people tend to over-insure. The requirement “replicate the data lake to the second site” arrives, and nobody asks why.
The premise we usually start from: the raw layer doesn’t need duplicating. It’s recoverable from source. If you have change-data-capture set up against production databases, you can replay the history — it’s slow, but not fatal, and costs incomparably less than a second storage cluster. At today’s flash prices, the difference is especially stark.
But you can’t turn this into a blanket rule, because the platform is made up of layers with very different costs of loss.
Raw layer and historical depth — not replicated, restored from source. Data marts, finished datasets, and customer extracts — exported to object storage; the volume is moderate, the value is high, and recomputing them from scratch takes too long.
The online feature layer — this is where the common mistake happens. If real-time scoring or fraud detection is built on top of analytics, the feature store for online lookups typically needs around 30 ms read latency at the 95th percentile and availability of 99.95%+, because its unavailability stops the entire business process. The volume, though, is small — tens, rarely hundreds of gigabytes. It should be fully replicated, and that’s cheap.
Metadata, the catalog, schemas, and pipeline code — the volume is negligible, and without them, restored datasets turn into a pile of files nobody remembers the contents of. Mandatory to replicate.
Models and the decision log — mandatory to replicate, and not for a technical reason this time. We’ll come back to that in the next section.
Circuit four · Artificial intelligence
Between renting and staying local
One boundary is economic and is set by utilization. The other is a latency boundary, and no amount of money moves it.
This is where the rent-versus-own argument gets sharpest, because accelerators are expensive, scarce, and age faster than a tender can run its course.
Let’s start with the economics, and it comes down to one number: sustained utilization. The rule the market keeps repeating: below roughly 40% load, renting wins; above it, owning wins, and round-the-clock inference makes owning the only real option. The spread in rental rates, meanwhile, is enormous: a survey of 25 providers found a nearly 14x difference for the same card. So both “renting is cheaper” and “renting is more expensive” are equally provable claims — it all comes down to the choice of provider and the load profile.
A natural split follows from this. Training and retraining models — an episodic, spiky workload — we rent. Production inference running round the clock at a steady rate — owned hardware. Experimental and side projects — rented almost always, because their horizon is shorter than even the first operating cycle, let alone the second.
Now the legal side, and over 2026 it changed twice, in opposite directions.
Data localization got easier. Law No. ZRU-1125 of March 26, 2026, restated Article 27-1 and repealed the previous blanket requirement to store all personal data inside the country. Mandatory localization now applies to three categories: biometric data, genetic data, and data of telecom operators’ users. For all other personal data, processing outside the country is allowed provided one of the three conditions in the same article is met.
Keep an important consequence in mind: voice is biometric data. Call-center recordings and voice prints fall under mandatory localization, and no amount of masking changes that.
Voice is biometric data. No amount of masking changes that.
It got stricter on the actual use of models. A law dated January 21, 2026, established that legally significant decisions affecting a person’s rights and freedoms may not rely solely on the output of AI-based systems. Unlawful processing of personal data using such technologies now carries a fine of 50 to 100 base calculation units, with confiscation of the means used to commit the violation.
For scoring, this is a direct architectural requirement, not a legal formality. You need a human in the decision loop, you need explainability, and you need a log: which features went in, which model version ran, which employee signed off on the result. All of that lives with you, not with your compute provider — which conveniently lines up with good engineering practice anyway.
As for masking data before sending it to an external environment — it’s a workable technique, but its limits are narrow. For tabular features, deterministic tokenization works well: relationships between tables are preserved and model quality barely suffers. For free text, and especially for voice, there’s no reliable masking: the entity recognizer makes mistakes, and a person can be re-identified from the combination of remaining features. So “mask and send” only holds up for structured features.
One last point on this circuit. Owned, local hardware for generative workloads is needed wherever response speed matters. A voice bot in a call center is exactly that case: a person in a conversation expects a reply within roughly 800 ms, the median pause between turns in live speech is around 200 ms, and past 1.5 seconds the caller assumes the line dropped. Within that 800 ms budget you have to fit end-of-utterance detection, final recognition, first-token generation, first-chunk speech synthesis, and the network round trip both ways. An external cloud eats that budget before the model even starts speaking.
Choosing where to rent
Where exactly to rent: nearby or far away
The choice hasn’t come down to “own hardware or a global hyperscaler” for a long time now. A local option covers more scenarios than people tend to assume.
The rent conversation is usually framed as a choice between owned hardware and a global hyperscaler. In our market, that hasn’t been true for a while, and the local option deserves its own look.
There are domestic operators with their own Tier III data centers, IaaS, object storage, and backup as a service. Compared to a global cloud, what they offer is: single-digit millisecond latency instead of tens of milliseconds, domestic jurisdiction, a contract and invoice in local currency, and a live engineer you can actually go visit. For the three data categories under mandatory localization, this isn’t an advantage — it’s the only lawful rental option.
For accelerators, the situation has changed faster than commonly assumed. There are already several domestic operators offering them for rent, the cards were bought along with the servers, and at some providers that capacity sits noticeably underutilized. The shortage reported for the global market doesn’t carry over well to local rental right now: domestic demand simply hasn’t caught up with supply. So the thing to check isn’t whether the cards exist at all, but the card generation, memory per card, and the contract terms.
What the local market usually still lacks: depth of service catalog, and mature managed databases and queues. That’s rarely critical for the four circuits in this piece, but it’s noticeable if you were counting on offloading platform operations to the provider.
A separate point that almost always gets left out of rental calculations: there’s no such thing as pure renting. At the boundary with a rented site, you’ll almost certainly need your own firewall — a core or NGFW-class one — because data flows both ways, tunnels need to terminate somewhere, and nobody’s going to write your perimeter policy for you. The switch you need may not be in the provider’s catalog either, and you’ll have to bring your own there too. These are capital costs, they show up in month one, and in any rent-versus-buy comparison they belong squarely in the rent column.
Global clouds win where you need a rare service, an order of magnitude of elasticity, or capacity that simply doesn’t exist in the country. And they lose on everything else: latency, cross-border transfer, egress cost, support in your own time zone. But the main point for us is the outbound link. Uzbekistan’s external connectivity is limited both in capacity and in the number of independent routes, so once anything significant moves to a global cloud, the resilience of that external link becomes part of the whole system’s risk model — not a line item in a contract with a connectivity provider. It needs to be accounted for in the same place you account for site failure.
For the four circuits covered in this piece, local rental covers more scenarios than commonly assumed: the object layer for backup copies, the landing zone, side projects, test environments.
Checklist
What to ask before pricing the spec
11 questions after which the rent-versus-own argument stops being a matter of taste.
Summary
In place of a conclusion
The rent-or-own argument almost always comes down to three numbers: the distance between sites, sustained utilization, and the project horizon relative to the 3-year first cycle and 3-year second cycle. Everything else is implementation detail.
Distance determines whether a synchronous scheme is even possible, and no amount of money buys that. Utilization determines which side the economics favor: below 40%, renting wins almost every time; with steady round-the-clock load, owning almost always wins. The horizon determines whether a purchase has time to pay for itself before the hardware or the licensing model shifts underneath you.
And on top of these three numbers now sits a fourth: the price of memory. As long as hardware’s share of the estimate holds around 80%, any argument about saving on licenses only touches one-fifth of the budget. You need to calculate where the money actually is.
Our working answer across the four circuits is this. Core virtualization — owned, at least 2 sites, but the standby doesn’t have to be full or the same generation as the primary. Backup — a full complex on the first site, media servers, rented object storage, and a landing zone on the second. Analytics — owned hardware in one center, outputs and datasets in object storage, with the online feature layer and metadata fully replicated since they’re small. Artificial intelligence — training rented, production inference and anything involving voice kept local.
If you’re facing a similar fork in the road and want to check the numbers before they go into a tender — get in touch, we’ll work through it together. Arguing with numbers is always more productive than arguing with slide decks.