Discuss a project

Issue 11Data infrastructure

Layer by layer

Continuing the conversation about kilowatts. Last time we calculated from specifications — from what companies were about to buy. This time we calculate from what's already installed: two sites, five generations of equipment, engineered systems with their own rack, and two temperature investigations — in one of which the obvious explanation turned out to be wrong.

Last time we calculated kilowatts from specifications — from what companies were about to buy. This time we calculate from what’s already installed.

The difference between these two calculations is bigger than it looks. A specification describes a single order: this many nodes, this configuration, one generation, one purchase. A rack doesn’t hold an order. A rack holds history: what was bought six years ago, what was added three years ago, and what arrived last month. Each layer came out of its own budget, with its own justification, and nobody added up the total, because the total never appeared in any single document.

What follows is a breakdown of two sites from different projects. One has three generations of hyperconverged nodes plus someone else’s equipment in the same rack; the other has a fleet that has been growing since 2012. Plus a separate section on engineered systems that arrive in their own rack, where the rules are entirely different. Power consumption is calculated from the equipment list — it’s an estimate. Temperatures are measured on running machines.

As last time: not all of these projects were built by us — in some we were involved at specific stages only. We don’t name the customers, and we don’t specify the industry or country.

What sits in one mixed rackThree generations of hyperconverged nodes, a previous-generation array, and switches. Not a singleaccelerator.11.25 kWCurrent-generation nodes, 9 pcs4.20 kWPrevious-generation nodes, 5 pcs3.40 kWNodes two generations back, 5pcs1.75 kWArray controllers and shelf0.56 kWSwitches04812 kWWhole rackUnits occupied33 of 42Power draw21.2 kWCurrent per phase30.7 Aa 32 A feed gives 25.6 AusableCores832Memory26 TBCapacity570 TBEstimate based on the equipment list.What sits in one mixed rackThree generations of hyperconverged nodes, aprevious-generation array, and switches. Not a singleaccelerator.Current-generation nodes, 9 pcs11.25 kWPrevious-generation nodes, 5 pcs4.20 kWNodes two generations back, 5 pcs3.40 kWArray controllers and shelf1.75 kWSwitches0.56 kW04812 kWWhole rackUnits occupied33 of 42Power draw21.2 kWCurrent per phase30.7 Aa 32 A feed gives 25.6 A usableCores832Memory26 TBCapacity570 TBEstimate based on the equipment list.
Fig. 1. Composition and power consumption of one mixed rack with three generations of nodes

Dense nodes

Hyperconvergence: power and temperature

A node packed with drives draws twice the power of an ordinary server and gets worse airflow. And the monitoring thresholds are still set for the previous generation.

Let’s start with hyperconverged nodes, because they’re the most common type of machine at the sites we’re discussing, and because they have a feature that rarely gets mentioned.

An HCI node is a server packed full of drives. An ordinary database server has two to four drives; an HCI node has eight, twelve, or twenty-four, all solid-state. That has two consequences.

First, on power draw. Drives account for a fifth of a node’s consumption, and they draw it almost constantly: idle, they could pull half as much, but in a cluster, background operations, deduplication, and replication never let them get there. A 172-core node with two dozen drives draws around 1,780 W, against 695 W for a database server with processors of the same class.

Second, on heat, and this matters more. A densely packed 1U node with drives across the entire front panel leaves little room for airflow. Everything coming in the front passes through the drives first, then the memory, then the processors, and exits the back already well heated.

Temperatures of dense nodes installed with no gapsMeasured on a running cluster. There is headroom up to the limit the processor reports, but monitoringreports critical overheating.monitoring system threshold — 70 and 75°76°Processor, cluster maximumhigh limit 92° · critical 100°72°Processor, typical valuehigh limit 92° · critical 100°51°Memory, maximumvendor threshold is higher,monitoring threshold 50°51°Chipsetvendor threshold is higher,monitoring threshold 50°020406080100 °CThe nodes sit flush against each other, with no gaps between chassis. Default monitoring thresholds fire long before theones the processor itself reports.Temperatures of dense nodes installedwith no gapsMeasured on a running cluster. There is headroom up tothe limit the processor reports, but monitoring reportscritical overheating.monitoring system threshold — 70 and 75°Processor, cluster maximum76°high limit 92° · critical 100°Processor, typical value72°high limit 92° · critical 100°Memory, maximum51°vendor threshold is higher, monitoring threshold 50°Chipset51°vendor threshold is higher, monitoring threshold 50°020406080100 °CThe nodes sit flush against each other, with no gaps betweenchassis. Default monitoring thresholds fire long before the onesthe processor itself reports.
Fig. 2. Temperatures of hyperconverged nodes installed flush against each other, versus thresholds

Measurements from a running cluster show this directly. Current-generation nodes installed flush against each other, with no gaps, read up to 76 degrees on the processors. Memory — up to 51. Chipset — up to 51. The processor itself reports a normal-operation ceiling of 92 degrees and a critical threshold at 100, with actual clock throttling kicking in even higher. So there is headroom, but less than you’d want, and less than the same processors get in a less densely packed chassis.

This leads to a practical recommendation that costs nothing at install time and is almost never followed: don’t stack dense nodes in an unbroken column. A single empty unit between groups of nodes, closed off with a blanking panel, noticeably changes the airflow pattern. And if there are no spare units, it’s all the more important not to leave open gaps elsewhere in the rack, because hot air will find its way back through them.

And a separate point about monitoring thresholds. Standard monitoring templates typically set processor temperature thresholds at 70 degrees for a warning and 75 for critical. That made sense for previous-generation processors. For current ones, which themselves report a normal-operation ceiling of 92 degrees, it means a stream of false alarms: the system reports critical overheating around the clock on a machine that’s operating normally.

The consequences are worse than they look. The on-duty shift gets used to the color red and stops reacting to it. And when a real overheating event happens, it drowns in the same stream of alerts. Thresholds need to be brought in line with the values the specific processor model itself reports — a one-time job that takes half a day, and it gives monitoring its meaning back.

While you’re at it, check the thresholds on the other sensors too. In the same batch of alerts we came across a disk array response-time threshold set at one ten-thousandth of a microsecond — a value no hardware could ever produce. A threshold like that fires constantly and means nothing.

The on-duty shift gets used to the color red and stops reacting to it. A real overheating event will drown in the same stream.

Accumulation

Three generations in one rack

Twenty-one kilowatts in a rack with not a single accelerator in it. No individual purchase ever looked at the total.

The first site. The rack holds two groups of hyperconverged nodes from different generations, a disk array with an expansion shelf, two fabric switches, and a management switch.

The old group is ten nodes: five with 32 cores and five with 48, each with 768 GB of memory and hybrid storage. The new group is nine nodes with 48 cores each, 2 TB of memory, and fully solid-state storage on every node. Plus a previous-generation array with a 24-drive shelf.

Adding it up: 33 units and 21.2 kW under typical load. 832 cores, 26 TB of memory, 570 TB of logical capacity.

Twenty-one kilowatts in a rack with not a single accelerator. For comparison: the cluster of 15 modern nodes we discussed last time drew 18.2 kW and didn’t fit on a single feed. Here the number came out higher — using ordinary nodes, simply accumulated over three generations.

The per-phase picture is just as unpleasant: 30.7 A per phase against a working limit of 25.6 A for a 32 A feed. The rack fails on feed capacity. And it still has nine free units. This is exactly the case last issue was written for — except here it came together on its own, without anyone deciding it should, simply because every purchase looked at its own units and never at the total.

The second site belonging to the same customer is set up similarly: 29 units, 21.4 kW, four previous-generation servers in place of an array.

Efficiency

The metric depends on what’s scarce

By cores per kilowatt, the new generation loses to the old one. By terabytes, it wins three times over.

Last issue had a section on watts per core, where it turned out that a denser processor was more efficient. Here that logic flips, and it’s worth walking through why.

Three metrics, three different answersEfficiency isn’t measured “in general” but by the resource that is scarce.Cores per kilowatt4732-corenode5748-corenode38New48-corenodeGigabytes of memory perkilowatt112932-corenode91448-corenode1638New48-corenodeTerabytes of capacity perkilowatt1832-corenode1448-corenode40New48-corenodeThe same set of nodes: the new generation loses on cores and wins on memory and capacityThree metrics, three different answersEfficiency isn’t measured “in general” but by the resourcethat is scarce.Cores per kilowatt32-corenode4748-corenode57New 48-corenode38Gigabytes of memory per kilowatt32-corenode112948-corenode914New 48-corenode1638Terabytes of capacity per kilowatt32-corenode1848-corenode14New 48-corenode40The same set of nodes: the new generation loses oncores and wins on memory and capacity
Fig. 3. Three efficiency metrics on the same set of nodes, and three different answers

Let’s compare whole nodes, not just processors.

The old 32-core node delivers 47 cores per kilowatt. The old 48-core node delivers 57. The new node, also 48 cores, delivers only 38. By cores per kilowatt, the new generation loses to the old one by a factor of 1.5.

The reason isn’t the processors. The new node carries three times more memory and four times more capacity, and all of that draws power too. By terabytes per kilowatt, it wins by a factor of 2.3-2.8; by memory, by a factor of 1.5.

Which brings us to what we consider the main takeaway of this issue. There is no single efficiency metric — the metric depends on what’s scarce. If you’re bottlenecked on compute, count cores per kilowatt, and a dense processor with minimal supporting hardware wins. If you’re bottlenecked on capacity, count terabytes per kilowatt, and the answer flips. Calculating efficiency “in general” is meaningless: the exact same pair of nodes will give opposite results depending on what you divide by what.

And there’s a practical consequence for any conversation about a fleet refresh. The claim “the new generation is more efficient” is true only to the extent that the new node mirrors the old one’s configuration. If memory and capacity grew along with the generation — and they always do — the node will draw more power, and a one-for-one swap will increase consumption rather than reduce it.

Fleet age

Layers nobody ever added up

Four racks, five generations, 25 kilowatts. Half of it comes from a layer that stopped being the primary one a long time ago.

The second site is a different customer and a different story. Here the fleet grew not in three layers but in five, starting in 2012.

Four racks, 83 units of equipment, 25.5 kW under typical load. Let’s break it down by age.

What the site's power draw is made of, by ageHalf the consumption comes from a single-generation layer that stopped being the core long ago. Afifth goes to the network.Serversfrom 2012Servers from 2015Servers2016–2018Servers2025–2026Networkequipment8%51%11%10%19%2.1 kW6 pcs13.1 kW30 pcs2.9 kW7 pcs2.5 kW5 pcs4.8 kW24 pcsFour racks, 83 units of equipment, 25.5 kW in typical operationEstimate based on the equipment list.What the site's power draw is made of, byageHalf the consumption comes from a single-generationlayer that stopped being the core long ago. A fifth goesto the network.8%51%11%10%19%8%Servers from 20122.1 kW · 6 pcs51%Servers from 201513.1 kW · 30 pcs11%Servers 2016–20182.9 kW · 7 pcs10%Servers 2025–20262.5 kW · 5 pcs19%Network equipment4.8 kW · 24 pcsFour racks, 83 units of equipment, 25.5 kW in typicaloperationEstimate based on the equipment list.
Fig. 4. Site power consumption by equipment age, and networking's share

2012-vintage servers: six units, 2.1 kW, 8% of the load. 2015-vintage servers: thirty units, 13.1 kW, more than half of the site’s total consumption. 2016-2018 servers: seven units, 2.9 kW. 2025-2026 servers: five units, 2.5 kW, a tenth of the total.

And on a separate line: networking equipment — 24 devices and 4.8 kW, or 19% of the site’s power. Almost a fifth of the electricity goes to switches, half of which were manufactured more than ten years ago. In individual racks, networking’s share reaches as high as 29%.

We’re calling out this number deliberately, because it breaks the usual assumption. When people calculate consumption, they count servers. Networking usually shows up in the math as “oh, and the switches too.” Here the switches eat more than the entire new fleet combined, and yet each one individually looks negligible — 390 W, who bothers counting that.

Networking’s share is also distributed unevenly. In racks where core and access switches are concentrated, it reaches a quarter and even 29%; in a rack holding only servers, it drops to one percent. This is worth factoring into planning: a rack with networking equipment needs its own calculation, and moving servers into it just “because there’s room” may not actually be possible.

The second observation is about the middle of the list. Thirty servers of one generation account for half of all consumption. This isn’t the oldest equipment or the newest — it’s the layer that was once the primary one and then never went away. It doesn’t get decommissioned, because something is still running on it, and it doesn’t get modernized, because the budget goes to new purchases. Over ten years, that layer becomes the biggest line item in the power bill, and nobody discovers it until someone sits down and adds it all up.

The switches eat more than the entire new fleet combined. And each one individually looks negligible.

Engineered systems

When the rack arrives as a whole

There’s a class of systems where power has already been calculated by the vendor before you buy. And where the rack is fully occupied from day one.

So far we’ve been talking about racks that are assembled in-house: buy the servers, rack them, connect them. There’s another class of systems where the rack arrives as a single manufactured unit, and the whole conversation about power works fundamentally differently.

These are engineered systems — database machines, big-data processing appliances, purpose-built backup systems. Inside, it’s ordinary server hardware, just with different component power ratings and mix. Outside, it’s a finished rack that already has multi-phase power distribution units installed, internal networking already wired, and cooling already calculated.

A typical starting configuration for such a machine is two database servers and three storage servers. In 2U chassis, that’s around 10 units of equipment plus internal fabric switches. But the rack takes up its full footprint — all 42 units — and it’s engineered for full population.

This is the main difference from a self-assembled rack, and it cuts both ways.

The good side: expansion doesn’t require recalculating power. Nodes are added one at a time or in blocks, based on the capacity and performance you need, and the question “will it fit within the kilowatt budget” never comes up — the vendor accounted for it in advance, the distribution units are already there, and the ratings are sized for a fully populated rack. This makes life much easier over a multi-year horizon: you know for certain you can grow into it without a single conversation with the facilities team.

The flip side: floor space and power feed are committed to the full rack from day one. You buy a quarter of the capacity, but the data hall has to allocate the footprint of an entire rack, plus a feed sized for its maximum. For a hall with a limited power budget, that can end up costing more than it looks like at purchase time.

Manufacturers publish the power feed requirements for these systems openly, in site-planning documentation: it specifies power draw, heat output, and feed requirements. It’s worth reading before the purchase, not after — unlike ordinary servers, here everything has already been calculated for you and simply written down.

Mainframes are their own story. There, absolutely everything is fixed: rack composition, power, cooling, connectors. The system arrives as a single unit, connects according to its own scheme, and there is no configuration flexibility by definition. From a data-hall planning standpoint, this is the simplest case of all: there’s one number to prepare space for. From a flexibility standpoint, it’s the exact opposite of a self-assembled rack.

The practical takeaway: if your fleet has, or is planning to add, a system like this, you can’t fold it into the general rack arithmetic alongside ordinary servers. It’s a separate planning unit, with its own power feed, its own heat output, and its own growth ceiling known in advance. And this is probably the one case where the question “will the expansion fit” has an answer before the expansion is even needed.

Three ways to fill a rackIn an engineered system, power is calculated before purchase. The price is space and a feed for a fullcabinet from day one.Self-assembled rackCompositionyou choosePoweryou calculatePDUsyou match them to the loadExpansionrecalculated every timeTaken from day oneas much as you installedGrowth ceilingunknown in advanceEngineered systemCompositionfrom the vendor's cataloguePowercalculated in advancePDUsmulti-phase, already installedExpansionby nodes, no recalculationTaken from day onethe whole cabinetGrowth ceilingknown before purchaseMainframeCompositionfixedPowerfixedPDUsown schemeExpansionwithin the cabinetTaken from day onethe whole cabinetGrowth ceilingset by the modelThree ways to fill a rackIn an engineered system, power is calculated beforepurchase. The price is space and a feed for a full cabinetfrom day one.Self-assembled rackEngineered systemMainframeCompositionyou choosefrom the vendor's cataloguefixedPoweryou calculatecalculated in advancefixedPDUsyou match them to the loadmulti-phase, already installedown schemeExpansionrecalculated every timeby nodes, no recalculationwithin the cabinetTaken from day oneas much as you installedthe whole cabinetthe whole cabinetGrowth ceilingunknown in advanceknown before purchaseset by the model
Fig. 5. Self-assembled rack, engineered system, and mainframe: how power planning differs

Cooling

The limit of air

A handful of industry thresholds, past which the cooling scheme changes not by choice but by physics.

The power we’ve been calculating has to be removed as heat. For cooling, it’s enough to know a handful of thresholds, and they’re industry-wide, not vendor-specific.

Up to about 10 kW per rack, an ordinary hot-aisle/cold-aisle scheme is enough, provided the aisles are genuinely separated and spare units are closed off with blanking panels. In the 10-30 kW range, you need in-row cooling units and careful control over where the air actually goes. At 30-40 kW, air cooling hits a physical limit: the volume and velocity of air that needs to move through the rack exceeds what a standard system can deliver. A rear-door heat exchanger can stretch that limit somewhat. Above 40 kW, liquid cooling stops being an option and becomes a requirement.

How to cool a rack depending on its powerIndustry thresholds. The top racks of both sites have already gone past the point where aisles andblanking panels are enough.racks of the sitesreviewedAisles and blanking panelsstandard hot and cold aisleseparation0–10 kWIn-row coolingin-row coolers, control over airdistribution10–30 kWRear-door heat exchangerextends the air scheme to itsphysical limit30–40 kWLiquidcold plates on chips; arequirement, not an optionabove 40 kW0102030405060 kWHow to cool a rack depending on itspowerIndustry thresholds. The top racks of both sites havealready gone past the point where aisles and blankingpanels are enough.racks of the sites reviewedAisles and blanking panelsstandard hot and cold aisle separation0–10 kWIn-row coolingin-row coolers, control over air distribution10–30 kWRear-door heat exchangerextends the air scheme to its physical limit30–40 kWLiquidcold plates on chips; a requirement, not an optionabove 40 kW0102030405060 kW
Fig. 6. Rack cooling methods by power density, and industry thresholds

Both of the sites we’ve examined live in the 3 to 21 kW per rack range. That means the upper end has already moved out of the zone where aisles and blanking panels are enough, into the zone that needs careful air management. No one made a decision about that — another layer just got added.

It’s worth comparing the airflow you need against what the existing system can supply. A perforated raised-floor tile typically passes 300 to 600 cubic feet per minute. A 21 kW rack needs about 3,300 — that’s six to ten tiles, and all of that volume has to actually reach that specific rack rather than dissipate into its neighbors. That’s harder in a mixed rack than a uniform one: equipment from different vendors has different depths, different fans, and different clearance requirements.

A separate point about airflow direction. Servers almost always blow front-to-back, but switches come in two configurations, and it’s easy to end up putting a device in a mixed rack that pulls its intake air from the hot zone. In a uniform rack, this question never comes up; in a mixed one, it needs to be checked for every single item — especially if the equipment arrived from different shipments years apart.

Measurement

When it’s not about kilowatts

A case where the obvious explanation turned out to be wrong, and only a second measurement revealed it.

The last section is about a case where we were wrong, and that’s more useful than everything else here.

Four database servers, the most modest item in the entire fleet: 48 cores, about 470 W, 1U chassis. On inspection, they showed processor temperatures of 78 and 85 degrees against a normal-operation ceiling of 92 and a critical threshold of 100, and network card temperatures of 86 and 90 against a threshold of 95. That’s five degrees of headroom on a component that these breakdowns usually don’t even mention.

The obvious explanation was simple: the servers sat in the top of the rack, where cold-air supply is weaker. We moved them down and took a second measurement.

Processor temperature: 77 degrees instead of 78. Drives, three degrees cooler. Network card: exactly 90, same as before.

What’s more, on both relocated machines, all four network card ports read exactly 90.0. Not 89, not 91. With different airflow and different load, that shouldn’t happen — there should be at least a couple of degrees of spread. This isn’t the behavior of a sensor pegged by its environment; it’s the behavior of a controller holding a setpoint. The cooling management system considers 90 degrees acceptable, because the threshold is 95, and it doesn’t spin the fans any harder.

That also explains why relocating the servers didn’t help. The controller compensated for the cooler air by lowering fan speed and brought the temperature right back to the same setpoint. The conditions changed; the operating point didn’t.

The practical lesson here: if the temperature doesn’t change when conditions do, look in the settings, not in the data hall — the cooling profile in firmware, fan curves, thresholds in the management system. And separately, in the configuration. These servers have standard heatsinks and standard fans, whereas in our other projects, comparable processors were consistently specified with high-performance versions. That’s two lines in a configurator that look like an optional extra and barely register against the price of the server.

And one more detail visible in the same measurements. On both machines, the second processor runs 9-15 degrees cooler than the first. That’s not airflow — a gap that size isn’t explained by airflow. It’s a workload sitting predominantly on one memory node: half the machine is running under load, the other half is coasting. Temperature turned out to be an indicator of a problem that had nothing to do with cooling.

One more thing that usually stays out of the picture: server configurations are built against a “25-degree data center environment” parameter. That’s not a preference — it’s the condition under which the specified cooling actually works. If the air entering the rack is warmer than that, the margin is spent before the server is even switched on for the first time, and no amount of shuffling within the rack will get it back.

Criterion

When to retire old equipment

To failure, end of support, and lack of performance, it’s worth adding a fourth reason.

Let’s return to where we started — the layers.

The old node group from the first site can’t be expanded — it’s on a previous license type. It occupies 31 units across two sites, draws 15.4 kW, and carries non-critical and test workloads at around 40% processor utilization. Over a year, that’s roughly 135,000 kWh for the equipment alone, and about 215,000 including cooling.

Keeping it around as an elastic reserve is a reasonable decision: the hardware is paid for, it works, and there are workloads for it. But that decision now has a price, and it’s measurable. Fifteen kilowatts is more than the entire new cluster draws at one of the two sites.

Which brings us to a criterion that’s usually missing from the decommissioning conversation. Old hardware gets retired when it breaks, when support runs out, or when it can’t keep up on performance. We’d propose adding a fourth reason: when the kilowatts it occupies are worth more than the work it does. In a hall with a limited power feed, every old node takes up space a new one can’t have. And that has to be measured not in units, but in kilowatts and amps per phase.

Checklist

What to check at your own site

Ten actions that require neither an audit nor any instruments.

  1. Add up the power consumption of everything in the rack, including switches, arrays, and anything that arrived from someone else’s shipment. Not layer by layer — as a whole.
  2. Calculate the current per phase for each rack and compare it against the feed rating — for mixed racks, this is usually the bottleneck.
  3. Break down consumption by equipment age and see which layer is the most expensive. It’s usually neither the oldest nor the newest.
  4. Calculate networking equipment’s share separately. If it’s more than a tenth of the total, it’s worth checking what’s actually installed and how old it is.
  5. Check the temperature thresholds in the monitoring system and bring them in line with the values the specific processor model itself reports. While you’re there, check the other thresholds too: ones that always fire mean nothing.
  6. Check whether dense nodes are stacked in an unbroken column, and if possible, break them up with blanked-off gaps.
  7. Check the airflow direction on every non-server device in the rack.
  8. Take temperature readings from running machines, including network cards, memory, and drives — not just processors.
  9. Treat engineered systems as a separate planning unit, and take the power feed requirements from the manufacturer’s site-planning documentation.
  10. For every old node, answer the question: what matters more — the work it’s doing, or the kilowatts and units it’s occupying.

Conclusion

In place of a conclusion

A fleet is never bought all at once. It’s bought incrementally: one purchase at a time, one justification at a time, each time looking at the spare units and almost never at the total current per phase. A few years later, you end up with a rack nobody ever designed, whose actual characteristics aren’t described in any single document — because every document only ever described its own layer.

The good news is that anyone can add up the layers: it takes no audit and no instruments, just an equipment list and a calculator. The bad news is that the total usually turns out bigger than expected, and it’s discovered exactly at the moment when one more server needs to go in.

If you want to find out what’s actually in your racks, let’s calculate it together. While we’re at it, we’ll take temperature readings from what’s already running — in our experience, that’s where the more interesting findings tend to be, more than in any specification.

Discuss your project

Let’s discuss your project

Tell us about your platform or project — an engineer will reply on Telegram or by e-mail.

Message us on Telegram

Or message us on Telegram — the bot will pass your question to an engineer.