Rating agencies don’t grade last quarter’s profit — they grade the ability to service obligations. A company can be profitable and still get downgraded, because the structure of its obligations no longer matches its capabilities. Infrastructure has exactly the same kind of rating — except no one calculates it and no one publishes it. Obligations keep growing — the regulator, the business, the market — while the ability to service them is set by a fleet bought several years ago for entirely different requirements.
Over the past year and a half we’ve seen this in almost every project. And we’ve seen it in five distinct variants — not five degrees of the same neglect, but five different states a company can be in. One customer paid in advance. A second is expanding against a forecast. A third deliberately trimmed the project and wrote down what it was paying with. A fourth hasn’t made a decision in three months. A fifth assembled its infrastructure from whatever budget was available in each given year. Below we go through all five — with numbers, because without numbers this is just a conversation about feelings.
All cases are anonymized down to the region level. There are no names, no volumes from diagrams, and no sums here — and there won’t be.
How debt is formed
The bar went up, the fleet stayed the same
Almost every new requirement turned into a permanent line of spending, not a one-off project.
Let’s start with why debt forms at all, even when nobody makes a mistake.
Look at the requirements of 2020. The perimeter was closed off with a firewall without deep traffic inspection, and that was enough for everyone. A second site was a wish, not a commitment: people planned for it, talked about it, but didn’t always get there. Immutable backups were an exotic thing people talked about at conferences. Object storage inside a company was rarely seen at all — S3 was a word from public-cloud presentations. Virtualization was a single platform, in a cheap edition, with a perpetual license, and the question “what if” never crossed anyone’s mind. There were no accelerators in the budget at all, because there were no workloads that needed them. Regulatory requirements were softer, and the cost of downtime was lower.
Now look at the requirements of 2026. The perimeter needs inspection and application control. Two sites are a commitment, with proof that failover actually works. Immutable copies are a line item in the technical specification, not a wish. Object storage is a working tool — archives and backups are already written to it. Virtualization is a subscription whose term has to be reconciled with the version’s lifecycle. Accelerators are a separate budget line that needs kilowatts found for it. Regulatory requirements are stricter, and the cost of downtime is higher.
The most important thing in this picture isn’t that there are more requirements now. What matters is that almost every one of them turned into a permanent line of spending rather than a one-off project. A second site isn’t a purchase — it’s a second site forever: power, links, licenses, people, procedures. A subscription is a payment that never ends. Immutable copies are capacity you can’t repurpose along the way.
That gives us a working definition we’ll use from here on. Infrastructure debt is the gap between what the fleet can do today and what’s being demanded of it today. Interest on that debt isn’t paid in money but in constraints: the bigger the gap, the fewer viable next steps you have left. And infrastructure default is the moment when there are no viable steps left at all, and the only option remaining is to replace everything, all at once, and at a premium.
And a caveat right away, without which the rest won’t make sense. Debt on its own isn’t a vice. Every infrastructure carries some, because requirements change faster than hardware depreciates. The only question is whether you know your own debt, whether it’s written down, and whether it’s built into the plan.
Five states on the scale: advance payment, growth against a forecast, restructuring, a deferred decision, forced patchwork. One section for each, below.
State one
Advance: paying before you need it
Exiting a lease forced the company to build everything at once, in full. Every sixth node in the fleet carries no application load — and that’s deliberate.
The first customer was moving from a rented cloud onto its own hardware. The decision was made once, and in full, so the fleet was built once, and in full: three sites, more than 40 worker nodes, plus a separate 3-node management cluster at each site.
Those management nodes are the most telling part of this whole issue. Every sixth node in the fleet carries not a single application VM. It exists purely so the site can be managed independently of the rest: serviced, updated, taken offline, without depending on its neighbors. In a debt-driven scenario, this is the first thing people cut — management gets folded into the production load because it’s cheaper at the start — and then for years they can’t take a site into maintenance without touching production.
The second thing paid for in advance here is unused headroom. Load across the three sites is split roughly 65-17-17: one is working, two are nearly empty by design. At the same time, about 71% of electricity costs are fixed and don’t depend on whether a node is carrying any load. In other words, the advance is paid not only as capital at purchase time, but every month in the power bill.
The third observation is about physics. One of the clusters — a dozen and a half nodes — had to be split across two racks, not because there weren’t enough rack units, but because there wasn’t enough power feed. It’s the same limit that, in advanced cases, shows up as “the expansion won’t fit”: here they simply hit it at the design stage, while it could still be engineered around, rather than in year five, when it can’t be anymore.
Advance payment has its own failure scenario, and we’re not going to pretend it doesn’t. If growth doesn’t arrive on schedule, a site running at 17% utilization will cost money every month, and by year four or five the fleet will start aging out without ever reaching its planned load. An advance is justified exactly to the extent that the forecast behind it is justified.
But this customer has an answer to that objection too. In parallel, they built independent addressing with BGP multihoming, and that changes the whole way the sites work: the plan is to spread the load evenly across all three so that none of them is primary and all three are peers. At that point the 65-17-17 split stops being a picture of “one working, two waiting” and becomes a transitional state. That, incidentally, is the best answer to the question of why you’d pay for unused headroom: you pay for it not as insurance, but so the sites become interchangeable.
State two
Growth against a forecast: expansion with justification
A reserve that can absorb half the load turns into a pumpkin at exactly the moment you need it.
The second case is a company with a rapidly growing customer base. Order of magnitude: several hundred thousand customers in summer 2026, with roughly a fivefold increase planned over a year and a half. With a forecast like that, arguing about “whether to expand at all” is pointless — the only real argument is about how.
Originally the sites were asymmetric: the primary carried the load, the backup sat as a hot standby running at single-digit-percent utilization. Expansion was sized not by “top up the primary” but for parity — so the backup site could take on the primary’s entire load, not half of it. That’s more expensive, and it’s the right call: a reserve that can only absorb half the load turns into a pumpkin, at a real failure, at exactly the moment you need it.
There’s one more decision in this project we like, because it’s about debt, not about specifications: capacity was split into two deliveries — part now, the rest a few quarters later, sized to actual growth rather than to a nice-looking number in a spec sheet.
Now for what’s invisible in the hardware in this project. Virtualization licenses were bought in January 2025 on a 3-year subscription, meaning they’re paid through early 2028. Mainstream support for the platform branch deployed at the sites ends in the fall of 2027. In other words, the version goes out of support before the paid term runs out, and the move to the next branch will have to happen within that term — including hardware upgrades wherever the current gear no longer makes the compatibility list.
This, by the way, is the answer to the popular argument “let’s take a simpler edition, it’s cheaper.” A cheap edition stops existing faster than you manage to save money on it: entry-level variants get phased out of sale over time, new versions ship only as part of higher-tier bundles, and support for the current branch ends on its own calendar, one that doesn’t care when you bought it. The lifespan of your infrastructure is measured not by the contract’s end date, but by the version’s end-of-support date. Those are different dates, and the second one is usually closer.
State three
Restructuring: debt taken on deliberately
Data volume grew fivefold, the budget didn’t grow at all. The difference between restructuring and cutting the budget is whether the decision was written down.
The third case is a backup project that had to be scaled back during preparation. The volume of data to be protected grew roughly fivefold over the course of the planning, while the budget didn’t grow at all. A common situation, and it usually ends with cuts made on the fly: licenses get dropped, the second site gets dropped, tape gets dropped.
Here they did it differently. Licenses and backup servers at both sites were kept per the original spec — the logic of the solution wasn’t compromised. What got cut was something else: a previous-generation tape library moves to the backup site instead of a new purchase, an existing array with high-capacity disks gets connected as a cold tier instead of a new one, and fabric switches aren’t purchased at all because there are enough free ports at both sites already.
Here’s what that actually means. The customer took on debt deliberately: part of the backup infrastructure now runs on previous-generation hardware, which has its own lifespan and its own spare parts kit. But that debt is calculated, written down, and backed by an argument: the core systems already replicate between sites through their own means, outside the backup system, so the lack of a fast tier at the backup site doesn’t leave a hole — it creates a recovery-speed constraint, and that constraint is known in advance.
The difference between restructuring and an ordinary budget cut comes down to exactly this: in one case you know what you paid with, and you can revisit it a year later; in the other, a year later you’re surprised.
State four
The deferred decision: a hundred and fifty unsupported nodes
The fleet looks uniform in the inventory system and splits into three groups in operation. The decision hasn’t been made in three months.
The fourth case is the most unpleasant, because nothing is happening in it. The fleet was bought in 2022: about 150 nodes of a single generation on a single platform, spread across more than ten purchase line items. Support ended a few months ago. The decision — renew or replace — still hasn’t been made.
First, what’s actually lost. Without an active contract, you lose all hardware service support: you can’t open a ticket, you can’t get a replacement part. Not part of the functionality — all of it. That said, the window for post-warranty servicing is always open, meaning the door hasn’t slammed shut — you can go back at any point, and that’s an important detail, because it turns the situation from hopeless into merely deferred.
Now for the composition, which is this case’s main discovery. The fleet looks uniform: one platform, one year, one vendor. In the inventory system, that’s one or two lines. In operation, it’s three different groups by purpose:
- about 90 nodes — hardware nodes for one virtualization platform;
- roughly 50 nodes — hardware nodes for a different virtualization platform;
- a dozen or so plain servers for standalone tasks.
Almost half of the entire fleet, moreover, comes from a single purchase line item.
Here’s where it gets interesting. Role determines the supporting gear, and the supporting gear determines everything else. Nodes on the hyperconverged platform are tied to its compatibility list: replacing a drive or a network card is decided not by what’s in stock, but by what the platform supports. Nodes in the second group are tied to both a compatibility list and their platform’s licensing calendar — two debts converge on them at once. The plain servers in the third group are tied to the lifecycle of the applications running on them, and nothing else.
The result is that the three groups have different replacement timelines and different triggers. And as long as you look at the fleet as a hundred and fifty nodes, the decision looks unmanageable: one large sum against a small recurring support payment, with no middle option in mind. Break it down by group, and three separate steps appear, each with a different timeline, each costed separately. It makes sense to start with the largest line item — it alone determines half the answer.
There’s also a second layer, invisible in the inventory records at all. The nodes were bought all at once, but they weren’t put into service all at once: the rollout stretched over more than 3 years, as need arose. Warranty and support, meanwhile, ran from the delivery date, meaning part of the paid term was spent on hardware that simply sat there waiting for its workload. Licenses were bought incrementally as nodes came online, and the license cycle was fixed at one year. The result is an overlay: the hardware ages from one date, the licenses live from several different ones, and over those same years the platform moved forward through versions while the last nodes were only just being brought online.
And one more consequence worth remembering on its own. The entire fleet was bought in a single year, which means the fans, drives, and power supplies are all wearing out in sync. Failures won’t arrive spread evenly across years — they’ll arrive as a window. Without an active contract, you meet that window with your own stockroom and your own hands.
State five
Forced patchwork: when every step was reasonable
Every step on its own was correct. The sum of the steps produced a structure with an active risk.
The fifth case started better than all the others. In 2020: a compact 5-node hyperconverged cluster, separate storage for backups, a tape library. A normal architecture for its time, with the priorities correctly set.
Then IT funding took a hit. From then on, every step was made out of whatever was available, and each one, on its own, was reasonable. More resources were needed — they built a second cluster, a dozen servers from a different vendor on a different virtualization platform, because that worked out cheaper. Capacity was short — they attached an array to it, even though the platform is designed for local disks. Databases lived on separate servers, bought separately, for the third time and for a third, unrelated, reason.
The result is three islands. Two hyperconverged platforms in two different network segments, plus database servers. And the link between the segments doesn’t run through a single device — it runs through a cascade of firewalls with no redundancy.
This is the one case out of five where the risk isn’t deferred but active. In all the previous ones, it’s “it’ll hit a wall in a year” or “it’ll get expensive in year four.” Here, any single device failing in the cascade breaks the link between segments today, and the longer the cascade, the more such devices it contains. And this isn’t the result of anyone’s incompetence: the non-redundant node exists precisely because the second segment was never planned at all — it emerged as a forced decision driven by money.
And separately — about the mirror-image relationship with the previous case. There, one platform covered three roles; here, three platforms cover one role. It looks like the opposite mistake, but the result is the same: a fleet that can be neither evolved in one move nor replaced in one move. Unifying the platform isn’t the same as unifying the fleet, and a variety of platforms isn’t flexibility.
Unifying the platform isn’t the same as unifying the fleet, and a variety of platforms isn’t flexibility.
Network debt
The network pays off its debt last
A switch just works — right up until the day you need to plug something new into it.
In these conversations, the network is the last thing anyone remembers, and that alone is a diagnosis. A server ages visibly: its support runs out, it drops off the compatibility list, it gets noisy and hot. A switch just works. It’s worked for ten years, it worked yesterday, it works today, and there’s no reason to touch it — right up until the day you need to plug something new into it.
Look at two data centers belonging to the same customer. The fabric is built from a mix of equipment: part of it discontinued and out of support from the manufacturer, part of it budget switches installed as a temporary fix. The fabric’s ceiling is 10G. The perimeter has no next-generation firewall. Management is manual, with no unified system. None of this bothered anyone for years, because the network worked.
Then, in that same year, the customer buys database servers and an AI server, and all of them have 25G network cards. There’s nowhere to plug them in. At that moment, network debt stops being an abstraction and turns into a number: the new server sits in the rack running four times slower than it could, or doesn’t run at all, until the fabric is replaced. And replacing the fabric isn’t like adding a node to a cluster — it’s a window into which everything falls at once.
A second example is smaller but more telling, because it’s about risk, not speed. An office network with its entire core running on a single 10G switch. There’s no spare of the same kind in the rotation pool: the pool consists of simpler gear, with fewer ports and interswitch links four times narrower. So if the core fails, the network will come back, but in a degraded state — and nobody has calculated that in advance.
Network debt has two features that make it more dangerous than server debt. First: the network pays off its debt last, because replacing it requires stopping everything, not just part of it. Second: network debt doesn’t show up in a report — it shows up in someone else’s project, when the server purchase is already done, the money already spent, and there’s nowhere to plug the hardware in. That’s why the list of support end dates for the fabric and the perimeter needs to be maintained exactly the way it is for servers, and reviewed before every server purchase, not after.
Metrics
What to measure debt with
Equipment age is a poor metric. What you should measure is the gap between what the fleet can do and what’s demanded of it.
Equipment age is a poor metric. A five-year-old node doing exactly what’s required of it creates no debt. A three-year-old one that fails to meet regulatory requirements does. Measure the gap, not the years.
The first measurement is the fleet’s age structure by consumption, not by count. At one customer’s site: four racks, more than 80 units of equipment, about 25 kW under typical load. Of that, servers from 2015 account for more than half of the site’s total power draw. Another roughly 19% of the load falls on network equipment, half of which was released more than ten years ago. That means the company’s production base is physically eleven years old, and any new project moves into what’s left over from it.
The second measurement is headroom calculated for a node failure, not for today. Take a cluster of 9 nodes, 508 virtual machines. Right now: 59.4% on CPU, 62.8% on memory, 55.7% on raw capacity. Looks calm. Recalculate the same thing for 8 nodes — that is, for the failure of one: 66.9%, 70.7%, and 62.6%. The calm ends there — memory headroom in N-1 mode drops below 30%, and the cluster will hit its limit on memory specifically.
And now for what turns this table from an observation into a deadline. The cluster started filling up rapidly in March. All of that load was accumulated in roughly six months, and at that pace, expansion is needed not in the next budget cycle, but now — by the time the hardware clears procurement, delivery, and deployment, there will be no failure headroom left at all.
The cluster filled up in six months. At the same pace, expansion is needed not in the next budget cycle, but now.
The third measurement is site symmetry relative to the load forecast, not relative to today. The question is simple: if the primary site disappears at the moment the customer base has doubled, will the backup take on all of the load, or half of it?
The fourth measurement is the support calendar against the planning horizon. The version’s end-of-support date, the hardware contract’s end date, the subscription term, and the depreciation term are four different dates, and they rarely coincide. Lay them out on a single timeline, and it becomes clear which one arrives first.
Default
From debt to technical default
What matters isn’t the size of the gap but the number of ways left to close it. Every deferred decision removes one.
Now let’s put it all on one timeline. On the day of purchase there’s no debt: the fleet matches exactly the requirements it was bought for. From there, requirements start rising in steps — a new regulation, a new system, a new perimeter-protection bar, a new platform version. Meanwhile the fleet’s capabilities don’t grow — they slowly decline: a version goes out of support, the hardware contract ends, power feed runs out, fabric ports run out.
Debt begins at the point where these two lines diverge. At that point it still breaks nothing and shows up nowhere in monitoring — it’s simply the moment from which a gap exists that needs to be closed somehow. From there, what matters isn’t how big the gap is, but how many ways you still have left to close it. At first there are many: extend support, add nodes, redistribute load, move part of it to the second site, delay by a quarter. Every missed decision removes one of them.
Technical default is the state where none of those ways are left. Expansion doesn’t fit within the power feed, a supported version for your branch is no longer sold, hardware support has to be restored through an inspection, the network can’t handle the new equipment’s speed, and recovery doesn’t fit within the required time. Each of these events is solvable on its own, and that’s exactly why each one gets deferred on its own. Together, they add up to a situation where the only option left is to replace everything at once, in full, at the highest possible price.
And this isn’t a hypothetical point somewhere in the future. We’ve seen it firsthand: a fleet that lived long enough to reach a point where switching application software was only possible through a full data export. Not a migration, not a move — a complete reload from scratch, because by that point not a single intermediate path was left.
Checklist
What to check in your own environment
Thirteen questions after which it becomes clear what state your fleet is really in.
- What new requirements has your infrastructure picked up over the last 5 years, and which of them became a permanent line of spending rather than a one-off project.
- How many nodes in your fleet carry no application load, and why — is it a deliberate advance or an accident.
- What share of site power consumption comes from equipment older than 7 years, and separately, from network equipment.
- What does cluster utilization come out to if one node fails, and which resource runs out first.
- How fast has the cluster filled up over the last six months, and at that pace, when will it hit N-1 mode.
- Will the backup site take on all of the primary’s load at the load forecast for 2 years out, rather than at today’s load.
- When does mainstream support end for the platform versions you have deployed, and does that line up with your subscription terms.
- Is there an active support contract for each equipment group, and exactly what do you lose where there isn’t one.
- How many groups by purpose does your fleet actually split into, if you look not at the model but at what determines a part’s replacement.
- What share of the fleet was bought in a single year, and when is synchronous wear-out of consumable components expected.
- What’s your fabric’s speed ceiling, and does it match the network cards on the servers you’re buying this year.
- How many non-redundant elements do you have on the paths between segments, and which of them appeared as a forced decision.
- Which of your budget cuts over the last 3 years are documented with a rationale, and which just happened.
Bottom line
In place of a conclusion
The five states in this issue aren’t a ranking of customers, and they’re not a scale from bad to good. They’re five ways of dealing with one problem: requirements grow faster than hardware depreciates. Each option has its own cost and its own failure scenario. There are really only two bad ones here, and both are recognizable by the same trait: nobody can say exactly what the company paid with.
Technical default doesn’t happen when hardware fails. Hardware always fails — there are spares, redundancy, and a maintenance schedule for that. Default happens when people stop making the decision — and one day discover there’s nothing left to choose from.
The good news is that, unlike a financial rating, you calculate your own infrastructure rating yourself. Once a year is enough: break the fleet down by group, and lay the dates out on one timeline. If, after that, you have three manageable steps instead of one unmanageable one, the debt is under control. If there’s one step left and it’s unmanageable — you already know what state you’re in. Want to walk through your own case — get in touch, and we’ll break it down along the same four measurements.