Discuss a project

Issue 06Data infrastructure

Four ages

How much support to buy up front, why keep spare parts under an active contract, how to tell that hardware can no longer be trusted one hundred percent, and when replacement turns out cheaper than renewal.

A procurement specification contains the line “warranty support — three years.” We ask where the three came from. The usual answer is that that’s what it was last time. This number determines how the customer will live through the next three years — and, if things go well, all ten — and it almost never gets calculated.

For us, the conversation about infrastructure age usually starts late — when something has already broken and a replacement part is on its way for who knows how long. By that point the choice is narrow, and every option is unpleasant.

Meanwhile, half the decisions are made at the entry point, on the day the specification is signed, and they cost nothing there. The other half come in year three and year six, when there’s time to think. Below we walk through every point where a decision can still be made calmly, along with the scale we use to sort hardware in a fleet into service classes.

Entry

What gets decided on procurement day

The support term and the scope of the first delivery are settled in one day and define the next three to six years.

The first decision is the support term. The practical answer: a base contract for three years, plus a clearly defined option to renew for another three. Not five, and not “however much they’ll give us,” but three plus three — because these are two separate decision points, and it pays to keep them apart.

The difference between “take five right away” and “three plus three” isn’t about money — it’s about information. Three years in, you know everything about the system: how it behaves, what has broken in it, whether its load is growing, whether the application software it was bought for is still alive. The renewal decision is made on facts. A five-year contract signed at the entry point is a decision made on a forecast, and a five-year forecast in infrastructure is a wish.

An important detail we covered in the first issue and will repeat here: the warranty clock starts at delivery, not at go-live. If the hardware sits in a warehouse for a month and then spends three months in commissioning and acceptance, you’ve already used up a third of a year.

And that’s the mild case. We regularly see another one: the hardware arrives, goes into storage, and sits there for a year, two, sometimes three — the project was postponed, the room isn’t ready, the application side isn’t done, the person in charge changed. All that time the hardware isn’t working, yet it’s aging on every front at once. The warranty is being used up, the generation is changing, and by the time it’s actually powered on, the system is no longer new: it’s a year or two closer to the second age than the books say. Add to that batteries and supercapacitors discharging on the shelf, and firmware several versions behind. If a pause between delivery and go-live is expected, it has to be raised at the procurement stage — by shifting the start of support, rather than discovering the loss later.

The second decision of the same day is the scope of the first delivery. A simple rule applies here: processors and memory go in at their target configuration right away; disks can be added later.

The reason isn’t price — it’s component lifespans. On most platforms, a processor can only be replaced together with the motherboard or within a strict compatibility matrix for that generation — so an upgrade two years later isn’t an upgrade, it’s a rebuild. Memory can formally be added, but it runs into the channel configuration: the new modules have to match the installed ones in type, capacity, and rank, or the system will either fail to build the configuration or build it with a loss of bandwidth. Disks are the friendliest in this respect — they’re added by shelves and groups, with no service interruption.

And right away, a caveat that has become more important than the recommendation itself over the past couple of years. Vendor cycles have gotten shorter. A component you planned to buy two years from now may by then be discontinued, replaced with a model of a different capacity, or moved to a different product line. Then the expansion turns into a mixed configuration with two drive types, different performance, and separate pools.

The practical takeaway: if expansion will definitely be needed, lock in not just the model but the ability to buy more of it — in writing, with a time horizon. If the vendor’s answer is vague, plan around a single delivery.

Lifecycle: where a decision can still be made calmlyFour points at which the choice is made on facts rather than under the pressure of a failure that hasalready happened.Procurementsupport term: threeyearsplus an option forthree moretarget CPU andmemory right awaydrives can be addedlaterYear threerenewal of the vendorcyclespare parts kitbecomes mandatoryfirst full diagnostichealth reportYear sixmaintain or replacecount four cost items, not justthe contract pricemove services down bycriticalityYear tenwithdraw from criticalworkloadsspare parts cannot bereplenishedthe decommissioningdecision is madebefore a failure, notafterFirst age0–3 yearsSecond age4–6 yearsThird age7–10 yearsFourth ageafter 10 yearsrisk on the vendorextended contractusually no supportpre-decommissioning03610yearsTrust in the hardware100%85–90%75–80%50%Lifecycle: where a decision can still bemade calmlyFour points at which the choice is made on facts ratherthan under the pressure of a failure that has alreadyhappened.Procurementsupport term: three yearsplus an option for three moretarget CPU and memory right awaydrives can be added laterFirst age0–3 yearsrisk on the vendorTrust in the hardware100%Year threerenewal of the vendor cyclespare parts kit becomes mandatoryfirst full diagnostic health reportSecond age4–6 yearsextended contractTrust in the hardware85–90%Year sixmaintain or replacecount four cost items, not just the contract pricemove services down by criticalityThird age7–10 yearsusually no supportTrust in the hardware75–80%Year tenwithdraw from critical workloadsspare parts cannot be replenishedthe decommissioning decision is made before afailure, not afterFourth ageafter 10 yearspre-decommissioningTrust in the hardware50%
Fig. 1. Hardware lifecycle and four decision points

Three plus three

Post-warranty: renew or let go

What gets decided in year three isn’t the fate of the hardware — it’s who pays for failures over the next three years.

The most common misconception about post-warranty support is that it just postpones the inevitable. “Let’s renew for another year and see.” A year is a bad unit: the administrative procurement procedure eats up months, and the next year it all repeats.

The scheme that works is different: after the base three years, you take one more full three-year vendor cycle. The fleet deliberately lives under contract for up to six years, and the question “maintain or decommission” comes up not in year three but in year six — when the hardware has genuinely reached that point.

Here’s what you’re actually buying with the contract, beyond replacing broken parts.

Access to updates. This is the thing people remember last, and it hits first. When support ends, access to firmware, microcode, and fixes closes. The hardware keeps running, but it stays on whatever version it had when the contract ended — with every known defect and vulnerability. For a system subject to regulatory requirements, that’s a problem in its own right, separate from hardware reliability.

The right to escalate. While the contract is alive, you have a procedure: a ticket, a priority, a response time. After it ends, you can buy one-off service on demand — at a different price, and with whatever turnaround you get.

A known price. Renewal costs around 10–15% of the price of equivalent new hardware per year. This is probably the one number in the entire issue worth memorizing, because everything else is calculated around it. Three years of renewal already amounts to a third to a half of the replacement cost. The conversation about whether a six-year-old system is worth maintaining starts right here and continues in Section Eight.

Let us also say what this scheme doesn’t include. There’s no “let’s live without support, and if something breaks, we’ll buy the part” option. It exists as a fact, but not as a decision: lead times for parts for post-warranty hardware are measured in weeks, and all that time the criticality of the system running on it doesn’t go anywhere.

Spare parts

Do you need spares under active support

The answer is yes. And the stock grows over the years — but it’s the range of items that grows, not the quantity.

Let’s start with the question we’re asked more often than any other: why keep a stock of spare parts if you’ve bought a contract with a guaranteed response time?

Because the response time in the contract and the time it takes for a part to appear on site are two different things. As a rule, manufacturers don’t keep a local spare-parts warehouse in the country. The part ships from a regional hub, goes through customs clearance, and before shipping, through the supplier’s internal procedures for vetting the recipient and destination. Each of these steps is legitimate and predictable on its own, but together they add weeks on top of what the contract says. In our experience, actual delivery to Tashkent takes five to fifteen days — and that’s when things are calm.

An example we already wrote about in the second issue: the robot in a modular tape library stopped. The replacement procedure itself took a day. Waiting for the unit took a month and a half, and for that whole month and a half the second copy simply wasn’t being created. And there was a contract in place.

That’s why a spare parts stock isn’t an alternative to support — it’s a complement to it. The contract covers money and the warranty; the stock covers time.

The contract covers money and the warranty; the stock covers time.

What goes into it

It makes sense to split the stock into two tiers.

The first tier is wear items. Things whose failure within a one-year horizon is likely, not hypothetical. Drives, power supplies, fan modules, controller supercapacitors, motherboard batteries, memory modules. The working norm for drives is 3–5% of the installed population, rounded up, with a minimum of two units. On a system with seventy-two disks, that’s four disks — and it isn’t overcaution, it’s statistics for a fleet of that age.

The second tier is active components. Controllers, adapters, transceivers, backplanes, and riser boards. With a live contract, these items are redundant. Without a contract, they’re what determines whether a node goes down for a week.

System boards are a separate line. They usually aren’t stocked because of their price, and that’s a reasonable decision — but it has to be made consciously and written down, because its cost is known in advance: no board means losing the entire node until a delivery arrives.

Four details people trip over

Only platform-specific parts. An appliance is built on a standard server platform, but it uses its own components with its own firmware. A drive from a regular server of the same model won’t work in such a system — it simply won’t be accepted. Orders have to follow the component list for the specific product, not the platform name.

The encryption variant. Some drives have separate part numbers without hardware encryption — they exist specifically for import into countries that restrict cryptographic equipment. You need to order the variant that’s already installed: otherwise you end up with a disk group made of drives in different variants, and that’s no longer a logistics problem but a configuration one.

Flash is ordered based on actual wear, not by norm. The remaining endurance of flash drives after five or six years is a measurable value, read with standard tools. If on some node it has dropped below ten percent, two cards in stock won’t solve anything: what’s needed there is planned replacement, not spares. This is the only item where the stock is calculated not from the population but from the readings.

The stock degrades by itself. Drives and supercapacitors need rotation; batteries self-discharge on the shelf. Once a year, stocked items should be run through a test bench. Spares that have sat in a cabinet for five years without ever being checked aren’t a stock — they’re an assumption about a stock.

Why the range grows

The key thing to understand here is that it’s the list of items that grows, not the quantity of each.

For the first three years, the stock can be minimal: everything is covered by warranty, and the only issue is speed. We keep what enables immediate service recovery — a couple of drives and a power supply.

Years four through six. Components with a limited service life start to fail. Controller supercapacitors last five to seven years, and on a six-year-old system they’re a top-priority item, not something exotic. Add fans and power supplies: mechanical parts and power electronics wear out predictably.

Year seven and beyond, especially without a contract, active components join the list: controllers, adapters, transceivers. Not because they’ve started failing more often, but because now there’s nowhere to get them quickly.

After ten years, the conversation shifts: the question is no longer what to put in stock but whether these items can be bought at all. If the platform has been discontinued, the spare parts stock becomes a non-renewable resource — you’re spending what can’t be replenished.

And one last observation that shows well where the real risk lies. On previous-generation arrays, flash drives served as cache: ten to twelve units for the whole system, configured as mirrored pairs. The most write-heavy component in the array was also the least numerous, and its stock had to be calculated not as a percentage of the population but by wear.

How the spare parts range growsIt is the list that grows, not the quantity of each item: the less support, the more there is that cannot beobtained quickly.0–3 yearsunder warranty4–6 yearsextendedcontract7–10 yearsusually nocontractafter 10 yearspre-decommissioningWear partsData drivesPower suppliesFan modulesController supercapacitorsSystem board batteriesCache flash componentsActive componentsMemory modulesControllers and adaptersTransceivers and cablesBackplanes and riser boardsItems decided case by caseSystem boardsShelf I/O modulesKeep in stockAs neededNot neededHow the spare parts range growsIt is the list that grows, not the quantity of each item: theless support, the more there is that cannot be obtainedquickly.after 10 years · pre-decommissioning7–10 years · usually no contract4–6 years · extended contract0–3 years · under warrantyWear partsData drivesPower suppliesFan modulesControllersupercapacitorsSystem board batteriesCache flashcomponentsActive componentsMemory modulesControllers andadaptersTransceivers andcablesBackplanes and riserboardsItems decided case by caseSystem boardsShelf I/O modulesKeep in stockAs neededNot needed
Fig. 2. Growth of the spare parts range as the fleet ages

Areas of responsibility

Who services a complex system

The vendor is responsible for the product. Nobody will do the on-site work for you.

On complex systems, the question “whose support is it” is framed the wrong way. The right answer is both, and they cover different things.

The vendor covers the product: parts replacement, known defects, updates, product consulting. What the vendor doesn’t do isn’t stinginess — it’s the boundary of the service. They won’t come and sort out your cabling. They won’t assemble the rack. They won’t check that everything is plugged back where it was after the work is done. They won’t plan a maintenance window around your calendar. And they won’t decide for you in what order to bring components up.

A real example from recent practice. A two-node system with an external disk shelf, after relocation to a new site. The standard topology-check utility on both nodes reports that the shelf is powered off. Meanwhile the shelf is on, warmed up, and the link LEDs on its modules are green. The operating system sees two system drives and nothing else.

The cables hadn’t been clicked all the way home.

One could go on about diagnostic levels here, and it’s a useful discussion, because that’s exactly the path taken: first the vendor utility’s output, then the bare operating system, then querying the controller bypassing the kernel, then the driver log. Each level said something different, and only a step-by-step descent showed that the problem wasn’t where people were looking. But it all ended with someone walking up to the back of the rack and looking at the connectors with their own eyes.

That distance — from a line in the utility’s output to a physical connector — is the integrator’s zone. No contract covers it.

Two practical consequences for planning any work on a complex system. Auto-start of the cluster software is disabled before the work begins: otherwise, after reassembly, the platform will try to assemble the disk groups on an incomplete topology, and you’ll get not “it didn’t start” but metadata corruption. And the topology is checked before the application layer starts, not after. The order of checks at power-on matters more than neatness during disassembly: you can take things apart imperfectly and fix it, but you can’t start up in the wrong order.

And a third point, about people. An engineer with fresh experience of this exact procedure is as scarce a resource as a spare part, and needs to be planned for just as far in advance. Someone who did the same relocation three weeks ago knows where the procedure stumbles and doesn’t spend the maintenance window figuring it out.

Who is responsible for what on a complex systemThe vendor covers the product, the integrator the work around it, the customer the requirements. Noarea covers the other two.Vendor supportReplacement of failedcomponentsAccess to firmware andfixesKnown product defectsProduct consultationRight to escalate underthe SLAReplacement warrantyResponsible for the product.The response time in thecontract and the time for apart to arrive on site aredifferent things.Integrator supportOn-site work andcablingPlanning maintenancewindowsLayered diagnostics:utility, operating system,firmwareCapturing the referencestateRelocation,commissioning, trainingManaging the spareparts stockResponsible for what liesbetween products. Nocontract covers the distancefrom a line in a utility's outputto a physical connector.Stays with the customerClassifying services bycriticalityAcceptable downtimeThe renew-or-replacedecisionBudget horizonSite access for peopleFleet composition andinventoryInfrastructure has no right totell the business how long itcan be down. It canhonestly say what eachoption will cost.Who is responsible for what on a complexsystemThe vendor covers the product, the integrator the workaround it, the customer the requirements. No area coversthe other two.Vendor supportReplacement of failed componentsAccess to firmware and fixesKnown product defectsProduct consultationRight to escalate under the SLAReplacement warrantyResponsible for the product. The response time in thecontract and the time for a part to arrive on site aredifferent things.Integrator supportOn-site work and cablingPlanning maintenance windowsLayered diagnostics: utility, operating system,firmwareCapturing the reference stateRelocation, commissioning, trainingManaging the spare parts stockResponsible for what lies between products. Nocontract covers the distance from a line in a utility'soutput to a physical connector.Stays with the customerClassifying services by criticalityAcceptable downtimeThe renew-or-replace decisionBudget horizonSite access for peopleFleet composition and inventoryInfrastructure has no right to tell the business howlong it can be down. It can honestly say what eachoption will cost.
Fig. 3. Areas of responsibility: vendor, integrator, customer

Everything ages

The whole infrastructure ages, not just servers

Fleet age is counted by servers and arrays. Yet switches, firewalls, batteries, and cables age too.

When people talk about obsolescence, they talk about servers and storage systems. Meanwhile, everything in the rack ages.

Network equipment becomes obsolete not because of its hardware but because of port speeds and software support. A switch can run for ten years without a single failure — and still become an irremovable bottleneck, because the new generation of connected devices is designed for a different bandwidth. We covered this in the second issue using a tape drive generation change as the example: the “switch replacement” line in the estimate looked like budget padding, but it was a direct consequence of the previous-generation port being unable to carry the new drive’s stream.

Firewalls age faster than anything else, for two reasons at once: inspection performance falls relative to the grown traffic, and software version support ends before the hardware wears out. Here age is measured not in years but in which features you can no longer turn on without losing throughput.

Uninterruptible power supplies are batteries with a three-to-five-year lifespan inside an enclosure with a ten-year lifespan. They need to be treated as a separate product with its own replacement cycle.

Cabling is the area where people cut corners most often. Multimode fiber is replaced with new fiber during a move, no discussion: microbends accumulate over the years, and adhesive and plastic dry out in the hot aisle. The problem is also that such a patch cord doesn’t fail honestly and all at once — it produces a growing number of errors under load, and the cause will be sought everywhere except in it. Ten- and twenty-five-gigabit copper connections require inspecting every single unit: they’re thinner, kink more easily, and their latches get stretched.

A separate note on AI accelerators — today the fastest-aging class of hardware in the rack. Generations change every year to year and a half, and this already directly affects procurement procedures: we’ve seen technical specifications for government tenders where the cards and configurations described were already the previous generation at the time of publication and simply couldn’t be ordered — approval took longer than the generation lived.

Hence a practical consideration worth raising before procurement. If an AI task is large and one-off — training, an experiment, testing a hypothesis — take the rented-capacity option seriously. Accelerators cost an astronomical amount, they change constantly, and when you buy, that whole cycle lands in the customer’s pocket along with the spare parts strategy: how to stock spares for a product that won’t be on the price list a year from now is a question without a good answer. Owning hardware is justified where the load is constant, the data can’t leave the premises, or the payback has been calculated over a horizon shorter than a generation change.

And the general rule that follows: the younger the component class, the shorter its obsolescence cycle. An array on spinning drives calmly ran for ten years, because nothing around it changed quickly either. Flash generations change twice as fast — not because flash is worse, but because the interfaces, density, and protocols around it move every two to three years. Hardware bought in a fast-changing class becomes obsolete relative to its environment before it physically wears out.

Four ages

The trust scale and service class

Hardware doesn’t go from “new” to “broken” in one move. There are four periods between those states.

This is what the whole issue was written for. The scale is simple; we use it for planning and in conversations with customers.

The first age, up to three years. Full trust. The wording matters here: not “the hardware doesn’t let you down,” but the risk lies entirely with the vendor. First-year failures don’t go away — infant mortality in electronics is real — but it isn’t the customer who pays for them. The scale doesn’t measure hardware reliability; it measures who bears the consequences of a failure. During this period, you can put anything on the hardware, including the most critical services, with no additional caveats.

The second age, four to six years. Trust 85–90%. The hardware is under an extended contract, but components with a limited service life start to fail. This is where the spare parts stock appears as a mandatory line item, not a wish. Critical services can stay, but they need a tested switchover mechanism, not just an entry in the disaster recovery plan.

The third age, seven to ten years. Trust 75–80%. The contract has most likely ended or now costs unreasonable money. The range of spares has grown to include active components. Test and pre-production environments, supporting services, and archive storage tiers live comfortably on such hardware. But you can’t put a backup site on aging hardware: a replica by definition must be no worse than what it protects, or the switchover will happen exactly once — and to the wrong place. Primary production — only if there’s a working replica on something younger behind it.

The fourth age, after ten years. Pre-decommissioning, trust 50%. Half is an honest estimate of what you’ll get at the moment you need the hardware. Nothing should remain on it whose failure you’d have to explain. It’s a resource for things you can afford to lose: lab benches, one-off tasks, a training ground for the team.

Next, the scale is combined with service criticality, producing a matrix that the entire fleet is distributed by. The rule that follows from it is simple: a service must not run on hardware whose trust level is lower than that service’s requirements. If a system must be back within four hours but runs on third-age hardware with no spare parts, the target exists only on paper — and that’s visible on paper too, without any outage.

Which service at which ageA service must not run on hardware whose trust level is lower than that service's requirements.First age0–3 years100%Second age4–6 years85–90%Third age7–10 years75–80%Fourth ageafter 10 years50%Core banking system, processingProduction databases under loadProduction virtual farmBackup site and replicasOperational backup tierDepartmental file servicesTest and pre-productionenvironmentsArchive and long-term retentionLab benches, trainingAllowedAllowed with caveatsWe don't place itWhich service at which ageA service must not run on hardware whose trust level islower than that service's requirements.Fourth age · after 10 years · 50%Third age · 7–10 years · 75–80%Second age · 4–6 years · 85–90%First age · 0–3 years · 100%Core banking system,processingProduction databasesunder loadProduction virtual farmBackup site and replicasOperational backup tierDepartmental fileservicesTest and pre-productionenvironmentsArchive and long-termretentionLab benches, trainingAllowedAllowed with caveatsWe don't place it
Fig. 4. Mapping service classes to the four ages of hardware

Signs of the end

When it’s time to decommission

Letting hardware go on time is a skill of its own. The skill is not waiting for a failure to serve as the reason.

Decommissioning almost always gets postponed, because hey, it works. A formal reason appears at the moment of a crash, and that’s the most expensive reason possible. Below are signs, any one of which is enough to start the conversation.

The platform has hit its ceiling. Not in performance, but in design limits: no more drives fit, no more memory is supported, cache capacity is limited in software and can’t be increased for any money. If the next expansion is impossible in principle, the hardware is already dictating your architecture rather than serving it.

The system’s composition is only approximately known. A real case: the composition of a storage system was reconstructed from photos of its shelves — how many shelves, what form factor, which drives inside. There was no documentation, and the people who had put it together were gone too. Until it’s clear what exactly is installed, there’s nothing to base a conversation about spares, support, or decommissioning on: there’s nothing to count.

Diagnostics show critical findings, and they don’t get resolved. A recent example: a full diagnostic report on a pair of systems gave them 92 and 93 points out of a hundred. Sounds great — until you read that within those hundred points there are three critical findings on each, fixes applied but never recorded in the fix log, uninstalled update components, a failed cache capacitor on one of the controllers, misconfigured non-volatile memory on all cells of one of the clusters, and nearly exhausted space in the system volume group. The overall score is an average across the whole hospital. What you need to read are the lines marked critical — and above all, look at how many of them have been hanging there for over a year.

Failures are no longer isolated. The first failure is covered by what was found in stock. The second — by what’s left. A spare parts shortage is discovered not at the moment of a failure but at the moment of the second failure, and that’s also the moment it becomes clear: repairing further means spending the irreplaceable.

Degradation without an obvious failure. The most unpleasant sign. Nothing is flashing red, but the error count grows, response times increase, and jobs stop fitting into their window. The hardware doesn’t fail — it stops coping, and this state can drag on for years until someone adds up the numbers.

And the last sign, covered separately in the next section: the application software no longer needs this platform.

Maintain or replace

When replacement is cheaper than support

Renewal costs 10–15% of the new price per year. Reliability isn’t the only thing to weigh against those percentages.

The “renew support or buy new” comparison usually boils down to a single line: the contract costs this much, the hardware costs that much, the contract is cheaper. And it really is cheaper — if you count only the contract.

Let’s run the numbers on two real generations roughly ten years apart. The task is the same: about 490 terabytes of usable capacity.

The previous-generation solution is a midrange array from the early 2010s, 3 TB drives, dual-parity protection in six-plus-two groups, plus hot spares. For 490 terabytes of usable capacity, that comes to 232 drives. At fifteen per shelf — sixteen 3U disk shelves, around 53 units together with the controller enclosure, which means two racks. The system does have flash, but as cache: ten to twelve 200-gigabyte drives in mirrored pairs, giving a software cache limit of 1–1.2 terabytes. That’s two tenths of a percent of usable capacity, and it can’t be increased for any money — the limit is set by the platform.

Today’s solution is a Lenovo ThinkSystem DG5200 with 24 NVMe drives of 30.72 TB each in a single shelf. After system reserves and protection — the same roughly 490 usable terabytes. Two units. Flash — one hundred percent of capacity.

The same capacity on two generationsRenewing support costs 10–15% of the price of new hardware per year. Reliability is not the only thingweighed against that percentage.Previous generation — midrange array, 3 TB drivesToday's solution — one NVMe shelf, 30.72 TB drives232 pcs24 pcsData drives53 U2 URack units occupied5 kW0.9 kWPower consumption10 W/TB1.8 W/TBWatts per usable terabyte12 pcs2 pcsSpare parts kit items0.2%100%Flash share of usable capacitySame usable capacity — about 490 TB. An estimate illustrating orders of magnitude: on a specific configuration thenumbers will differ, the ratio will not.The same capacity on two generationsRenewing support costs 10–15% of the price of newhardware per year. Reliability is not the only thingweighed against that percentage.Previous generation — midrange array, 3 TB drivesToday's solution — one NVMe shelf, 30.72 TB drives232 pcs24 pcsData drives53 U2 URack units occupied5 kW0.9 kWPower consumption10 W/TB1.8 W/TBWatts per usable terabyte12 pcs2 pcsSpare parts kit items0.2%100%Flash share of usable capacitySame usable capacity — about 490 TB. An estimate illustratingorders of magnitude: on a specific configuration the numberswill differ, the ratio will not.
Fig. 5. The same usable capacity on two hardware generations

Now four line items that don’t make it into the “contract versus hardware” comparison.

Power. A difference of about four kilowatts, continuous — roughly 36 megawatt-hours a year. And every kilowatt consumed also has to be removed as heat, so the bill ends up at around double that figure.

Space. Fifty-one freed-up units. This isn’t an abstraction: in the fourth issue we wrote that free units in a rack are the only cheap way to buy yourself the next upgrade without yet another move.

Spare parts. At the 3–5% of population norm, the old solution needs eight to twelve drives plus I/O modules and power supplies for sixteen shelves. The new one — the minimum two units and one shelf. That’s an order-of-magnitude difference, and that’s before the conversation about having to find parts for a discontinued platform somewhere.

Operational complexity. Two hundred thirty-two drives versus twenty-four isn’t just the probability of failure multiplied by ten. It’s sixteen shelves that have to be distributed across buses, racks, and power feeds, and two racks on the floor instead of two units.

Support stops being a saving: it defers the payment while leaving every other cost in place.

Let’s add it up. Three years of renewal — a third to a half of the replacement cost. Against that half stands a fivefold saving on power, a rack and a half freed up, a tenfold smaller spare parts stock, and moving off a platform that has hit its ceiling. With numbers like these, support stops being a saving: it defers the payment while leaving every other cost in place.

To be fair: this is an illustrative calculation of orders of magnitude, and on your configuration the numbers will differ. But the ratio — not percentages, but multiples — doesn’t depend on the configuration, and that’s exactly what should make it into the budget conversation.

And the most unexpected part

Everything above concerns the case where you make the decision. Often it’s made for you.

In our practice there was a fleet on a proprietary UNIX platform with previous-generation arrays: 146-gigabyte disks, a proprietary multipath manager, machine rooms with vertical heat extraction. It was replaced with classic x86 servers in 2021. The reason wasn’t wear, expired support, or a lack of spare parts. The application software changed, and with it went the need for that kind of architecture altogether.

The migration, meanwhile, turned out not to be a simple transfer. The data structure in the new system was so different that the standard tools didn’t fit: the database was exported to text files and loaded again from scratch. The data outlived the hardware, but it went through CSV.

Hence a conclusion worth keeping in mind for any replacement planning: the decommissioning date is set by the application system, not by the vendor’s datasheet. A plan built around the end-of-support date is planning the wrong thing. You need to ask not only “how much longer will the hardware last,” but also “what will be running on it three years from now, and will it even need this kind of platform.”

Checklist

What to ask about your fleet

Twelve questions that let you manage a fleet rather than react to it.

Twelve questions. If there’s a written answer to each one, you’re managing your fleet, not reacting to it.

Bottom line

Instead of a conclusion

Infrastructure age is neither a problem nor a diagnosis. The problem starts where age isn’t taken into account: when a ten-year-old system is trusted like a new one, when a recovery target is written for hardware that can’t meet it, when support renewal is bought out of inertia without a single calculation.

That’s exactly what the four-age scale is for — not to decommission everything in sight, but to make sure every system carries a load it can handle at its age. A seven-year-old array for archive is a sound engineering decision. The same array under the main database is a deferred outage.

And one last thing. We see many fleets that are kept running for years out of caution: replacing is scary, migration is a risk, and this way it works. The caution here is illusory. The risk doesn’t disappear because the decision was postponed — it just moves to a moment you don’t get to choose. Infrastructure needs to be assessed soberly, and you shouldn’t be afraid to expand and rebuild it — that’s cheaper, more predictable, and ultimately calmer than maintaining something you yourself no longer trust one hundred percent.

If you have hardware and it’s unclear what age it’s in, get in touch. We’ll assess its condition, place it on the scale, and calculate which is cheaper: maintaining or replacing.

Discuss your project

Let’s discuss your project

Tell us about your platform or project — an engineer will reply on Telegram or by e-mail.

Message us on Telegram

Or message us on Telegram — the bot will pass your question to an engineer.