Every technical specification we read describes the future infrastructure in cores, gigabytes, and terabytes. Sometimes it even gets down to port counts and patch cord lengths. We have never once seen a kilowatt.
Yet it’s kilowatts that decide whether what you bought will actually fit in the room. Units in the rack run out later — usually much later. What runs out first is the power feed, and right behind it, the room’s ability to remove the heat that power generates.
Here’s a telling story to start with. When designing one new data center, foreign architects allocated 3 kW per rack. That was a year or two before demand for model computation changed the market. Three kilowatts is four database servers, or two virtualization nodes. A single modern accelerator node wouldn’t fit into a rack like that at all, and an ordinary hyperconverged cluster would have to be spread across five racks. The project itself, though, was built to the standards of its time and contained no errors whatsoever.
What follows is a breakdown based on numbers from projects that have passed through our hands. One caveat up front: not all of them were built by us. Some we carried from calculation through implementation; in others we were involved only at certain stages — sizing the configuration, diagnosing a problem, preparing a proposal — while someone else did the actual work. That doesn’t make the specifications and measurements any less real, but we don’t want to take credit for work that isn’t ours.
We calculated consumption from the configuration bill of materials: we added up the rated power of the processors, the consumption of the memory modules and drives, the accelerators, network cards, fans, and platform, then divided by power supply efficiency. This is an estimate, not a meter reading, and we flag separately where it might diverge from reality.
This is the first part of the conversation: what consumption is made up of, how it depends on processors and accelerators, and how it’s distributed across a rack. The second part — about what’s already running at sites, about heat, and about turnkey complexes that arrive in their own cabinet — is in the next issue.
We don’t name customers: Central Asia, industry withheld, configurations generalized.
Factors
What a watt is made of
The calculation method depends on the type of machine. In one, processors decide everything; in another, they account for a twentieth of the total.
Before we get to racks, let’s break down what makes up the consumption of a single server. There aren’t many factors, but their weight varies enormously depending on the type of machine.
In a database server with two 24-core processors, almost all the consumption is the processors: 60% of roughly 700 W. Memory, drives, and networking together account for less than a fifth. A machine like that is easy to size: take the two TDP figures, add a third on top, and you get an answer with decent accuracy.
In a 172-core hyperconverged node, the picture is different. Processors contribute 39%, drives 21%, fans 14%, memory 13%. Sizing by processors alone here won’t work — the error runs one and a half times over.
And in a node with eight accelerators, processors take up 5% and barely matter. There, 69% is the accelerators themselves, and 9% is the fans, which have to push a huge volume of air through the chassis.
Hence the first practical rule: the calculation method depends on the type of machine. For ordinary servers, processors plus a margin are enough; for nodes with many drives, you need to count the drives; for accelerator machines, you can round off everything else.
And a few things that get left out of calculations most often.
Drives almost never rest. An enterprise NVMe drive draws 12 to 19 W under load, depending on capacity and type, and can drop to 4-8 W at deep idle. But in a hyperconverged cluster, idle almost never happens: background operations, deduplication, and replication keep the drives working constantly. So for a power budget it’s reasonable to take the upper bound — twenty-four drives in a node come to about 360 W.
Fans aren’t free. In a dense configuration they account for anywhere from a tenth to a seventh of total consumption, and their share grows as things heat up: the hotter the room, the more the server spends on cooling itself. That’s a feedback loop usually missing from calculations.
Power supplies have an efficiency rating. For modern ones it’s around 94-96% at typical load, but it drops off at both very low and peak load. Five percent of loss at the input is what the room has to supply on top of what the electronics actually consume.
Network cards are negligible in general-purpose servers, but in accelerator machines they become a noticeable line item: a few 400 Gb cards plus coprocessors add up to half a kilowatt in a node — as much as its processors.
Processors
Generations and efficiency
Per core, processors are getting cheaper to run. The server as a whole is getting more expensive.
Now for how all this changes over time — and here there are two opposing trends.
First: per core, processors are becoming more economical. Take processors from the same generation and calculate watts per core — a 96-core model comes out to 3.75 W, an 86-core one to 4.07, and a 24-core model from the same lineup to 10.62. The spread is almost threefold, and it runs against intuition: the denser the processor, the more economical each of its cores.
Let’s check this at scale, because at the complex level the difference looks even more impressive. A 45-node complex gives you 7,740 cores on 86-core processors. The same 7,740 cores on 24-core processors would require 162 nodes instead of 45. On processors alone, that’s 82 kW versus 31.5 kW — a difference of more than 50 kW, not counting memory, drives, and fans, which would also triple. Plus 162 units instead of 90, and twice as many switch ports.
The same holds across generations. A new generation, at the same core count, usually delivers more performance at the same or similar TDP — thanks to clock speed, cache, and memory bandwidth. If you choose a model by core count rather than its position in the lineup, that gain comes almost for free.
The second trend runs the opposite way: the server as a whole is getting hungrier. Top-tier models in the lineups have gone from 250 W per processor to 500 W. Memory got faster and there’s more of it. Drives per node went from four to two dozen. As a result, a new-generation node consumes more than an old one, even when its processors are more efficient per core.
These two trends don’t contradict each other; they just answer different questions. “How many watts per unit of work” is falling. “How many watts per rack unit” is rising. The room lives by the second number.
In the previous issue of this series, we wrote that under per-physical-core licensing, you should choose a processor by core count rather than by its position in the lineup. Here we add the other side of that same fork: density saves kilowatts and space, clock speed saves licenses. Which one is more expensive depends on which one you have less of.
Accelerators
Accelerator tiers
From one card in an ordinary server to eight in a dedicated chassis — the spread is wider than between all other machine types combined.
A separate word on accelerator machines, because the spread within this category is wider than between all the other server types.
The bottom tier is an ordinary two-processor server in a 2U chassis with a single 600 W professional card. A machine like that draws about 1,140 W — roughly twice a server without a card — and goes into any rack without special conditions. Over the past couple of years, well over a dozen such configurations have shown up in our region: it’s no longer exotic, but a working option for workloads that need a bit of acceleration.
The second tier is the same chassis with two cards: about 1,640 W. Still an ordinary rack, but already 820 W per unit — more than a hyperconverged node.
Then comes a break. Eight accelerators in an 8U chassis is a separate class of machine, and the generation-by-generation numbers look like this. The 2020-generation accelerator drew 400 W per chip. The next generation: 700 W. The one after that: 1,000 W. The current one: 1,100 W in air-cooled form and 1,400 W in liquid-cooled form.
In other words, over five years, power per accelerator has tripled, and an eight-accelerator node has gone from about 3.5 kW to 8.6 kW in the air-cooled version. The liquid-cooled version of the same generation delivers upward of 11 kW per node before processors and networking, and for it, direct liquid cooling isn’t optional — it’s mandatory: air can’t handle it.
This is worth pausing on, because this is where the boundary of applicability lies. As long as we’re talking about one or two cards, the question is settled in an ordinary room. Eight air-cooled accelerators can still be placed if the room has the power feed and airflow for it. The liquid-cooled version requires a loop that almost no one in the region has, and operating experience that’s rarer still.
A brief word on the next generation: it has been announced, but the manufacturer hasn’t disclosed public figures for power per accelerator or per rack, and the numbers circulating in reviews have no primary source. You can’t plan around them.
Over five years, the power of a single accelerator has tripled, and an eight-accelerator node has gone from three and a half kilowatts to eleven.
The practical takeaway is simple. If someone comes to you with “we need a bit of AI,” that’s most likely the first or second tier, and it gets solved in the existing room. If they come with a task at the level of model training, the conversation doesn’t start with accelerators — it starts with what the room actually has for power feed and cooling.
Placement
The rack doesn’t run out of units
A 15-node cluster takes up 30 units and doesn’t fit in a 42U rack. The bottleneck isn’t where people look for it.
Now for the arithmetic of placement, and here we have a real case.
A cluster of 15 nodes, each with two 86-core processors: 172 cores per node, 2,580 cores per cluster. It takes up 30 units — less than three-quarters of a rack. By typical consumption that’s 18.2 kW; by calculated maximum, 26.7 kW.
The cluster had to be split across two racks.
Not because there wasn’t enough space. There was space to spare: after the split, each rack has 16 and 14 units occupied out of 42 — more than half sits empty. What ran out was kilowatts. The customer paid for two rack spaces to house equipment that physically fits into one.
This, incidentally, is an indirect way to find out the real power feed in a room when there’s no nameplate data. Since 15 nodes didn’t fit together, the ceiling is below 18 kW. Since 8 nodes fit with margin to spare, it’s somewhere around 12-15 kW. That estimate is more accurate than what you usually get when you ask “how much do you have per rack.”
And here’s a consequence that wasn’t part of the placement task at all. After the split, each rack holds roughly half the cluster. That means a rack failure — a feed, a breaker, a switch — takes out half the nodes at once. In most quorum schemes, half is worse than a third: three racks of five nodes survive the loss of one; two racks of seven or eight is a question that has to be solved separately, and in advance.
The task was framed as “place the equipment.” The decision was made based on power. But the failure domain changed as a result, and nobody planned to recalculate fault tolerance, because that item was never on the agenda.
From this we made a rule that we now apply to every project: the rack placement plan needs to be coordinated with the platform architecture, not just with the installation crew. Otherwise the power constraint silently rewrites the reliability decision.
Distribution
Kilowatts arrive in phases
Fifteen kilowatts isn’t one number — it’s three phases, two feeds, and breaker ratings that all have to line up with each other.
Up to now we’ve been writing “15 kW per rack” as if it were a single figure. In reality, no such single figure exists — there are three phases, each with its own breaker, and two independent feeds. And how the load is distributed across them decides whether the setup works or trips the protection.
Let’s start with the arithmetic, without which nothing else makes sense. A three-phase feed delivers three times the power of a single-phase feed of the same rating: three phases at 230 V instead of one. The 1.732 coefficient that often comes up here belongs to the same formula written through the 400 V line voltage, and it’s not the right tool for comparing against a single-phase feed. But you can’t take everything printed on a breaker’s housing at face value: for continuous load, the accepted rule is 80%, and power distribution equipment manufacturers already build that into their nameplate ratings. Hence a table worth keeping in mind.
Now let’s come back to our 15-node cluster. Typical consumption is 18.2 kW, maximum 26.7 kW. A three-phase 32 A feed gives 17.7 kW of working power. So the cluster doesn’t clear the feed even on an ordinary day, let alone at maximum.
Let’s check it a different way, by phase. The ideal split is 5 nodes per phase, which is 6.05 kW and 26.3 A against a working limit of 25.6 A. It doesn’t clear even then, with flawless balancing. And if you split it carelessly — 6-5-4 — the first phase ends up at 31.6 A, which isn’t “tight” anymore, it’s a tripped breaker.
The split version looks different. Eight nodes give 14.2 kW at maximum and about 20.6 A per phase; seven nodes give 12.5 kW and 18.1 A. Both fit within a 32 A feed with a healthy margin. What looked like a “we ran out of kilowatts” solution turns out, on closer inspection, to be a “we ran out of feed rating” solution — and those are different things: the feed into the room could have been larger; what they actually hit was the capacity of whatever distributes that power across the rack.
Next: three mistakes we see most often in power calculations.
First: sizing by the power supply nameplates. A node with eight accelerators has eight 3,200 W power supplies, for a total installed rating of 25.6 kW, against a real maximum of about 12,600 W. A customer who puts the nameplate total into their request to the utility either gets turned down or pays for twice the power feed they actually need. Power supplies are sized with margin and for redundancy; their sum is not consumption.
Second, the opposite mistake and a more dangerous one: splitting the load between feeds. Redundancy via a scheme with two independent feeds means each feed has to hold the full load, not half of it. Anyone who assumes each feed carries its own share finds out about the mistake the first time one feed is taken down for planned maintenance — when everything shifts to the second one. Redundancy doesn’t halve the requirement on a feed; it doubles it: two feeds of 18 kW each, not two of nine.
Third: forgetting about the neutral and phase balancing. Server power supplies are a non-linear load, and when phases are unbalanced, current flows in the neutral that doesn’t cancel out. Hence the requirement to size the neutral with margin relative to the phase conductors and to watch for even distribution. In practice this means the order in which servers get plugged into outlets isn’t a matter of installer tidiness. On a vertical PDU, outlet groups are split by phase, and if you plug eight nodes in a row into adjacent outlets, they may all land on the same phase. The current on it will then run three times higher than calculated.
Hence a practice worth demanding from the installation crew and documenting formally: a wiring diagram showing which node goes into which outlet and onto which phase, with a current calculation for every phase of every feed — separately for normal operation and for the loss of one feed. PDUs with per-phase current metering cost more than ordinary ones, and that difference pays for itself the first time you need to figure out why the protection tripped.
And last, on uninterruptible power supplies. Their capacity is rated in kilovolt-amperes, not kilowatts, and older models have a power factor of 0.8. A 20 kVA unit in that case delivers 16 kW, not 20. Modern models have a power factor close to one, but that needs to be checked against the specific model, not assumed from the name. A mistake here shows up at the worst possible moment — when switching over to batteries.
Bill structure
What pays for simply being on
A 45-node complex across three sites. Seventy-one percent of electricity spend doesn’t depend on whether anything is actually working.
Next, the complex as a whole — and it reveals something that changes how you approach savings.
45 nodes across three sites, 7,740 cores, about 90 TB of memory. All three are working, but the load is distributed unevenly: the primary site carries about 65% of the work, the other two 17-18% each — that’s where auxiliary services live, along with the readiness to take over the primary workload.
In terms of consumption, that’s 17.8 kW at the primary site and 12.0 kW at each of the other two. 41.8 kW in total.
Now let’s split those 41.8 kW into two parts. An idle node isn’t a powered-off node: the memory, the two dozen drives, the network cards, the fans, and the platform keep drawing power regardless of whether the processor is computing. By our estimate, a node like that draws about 660 W against 1,210 W under load.
From this: 29.7 kW goes toward the equipment simply being on, and only 12.1 kW depends on actual work. Seventy-one percent of the electricity bill is paid for the fact of being switched on.
This breaks the usual optimization logic. Consolidating load, raising utilization, fighting for VM density — all of that works within the 29% that’s left. The levers that move the main part sit somewhere else, and they get pulled once: how many nodes you turned on, and how many watts a single core costs. After procurement, that part of the bill can’t be optimized at all.
Seventy-one percent of the electricity bill is paid for being switched on, not for computing.
The same calculation leads to an unexpectedly pleasant conclusion about standby capacity. If the standby sites did nothing at all, the complex would consume about 38.2 kW. An active standby costs 41.8 kW — a 9% addition to the bill. Yet roughly half again as much useful work gets done, because the auxiliary services compute on hardware that’s already paid for, powered, and cooled regardless.
In other words, the right framing isn’t “standby sits idle and that’s wasteful” — it’s “standby should be working, because you’re paying the bulk of its bill either way.” The only condition is that the workloads on the standby sites need to be ones you don’t mind displacing at failover. That’s a matter of policy, not electricity.
For power design, note, you still take the maximum. Every site has to be able to accept the full workload, which is about 27 kW versus 12 on an ordinary day. The gap between what a room must be able to withstand and what it actually carries day to day is more than double, and you need both numbers: the first for engineering, the second for the bill.
Checklist
What to check before procurement
Eight questions worth going through before the equipment arrives on site.
- What power feed per rack actually exists, not what the building’s design documentation says. If there’s no figure, calculate it from what’s already installed and running.
- How many kilowatts the equipment being procured draws in typical operation and at calculated maximum. These are two different numbers, and you need both.
- How power redundancy is set up: in a two-feed scheme, each feed has to hold the full load, not half.
- What the rating of the input breakers is, and what that gives you once you apply the 80% rule. Room power and the rack’s feed rating are two different constraints.
- Whether there’s a phase wiring diagram with current calculated for every phase of every feed — in normal operation and with one feed lost.
- What units the UPS capacity is rated in, and what its power factor is.
- How many racks are actually needed based on power, not units, and what that does to the platform’s failure domains.
- How many watts per core the chosen processor draws, and how that relates to the platform’s licensing metric.
Bottom line
In place of a conclusion
Infrastructure is purchased in cores and terabytes, but deployed in kilowatts. As long as these two figures live in different documents and get discussed by different people, the gap between them will keep surfacing at the installation stage — at the most inconvenient moment, once the equipment has already arrived.
The good news is that this can be calculated in advance, without any specialized knowledge. The configuration is known before procurement, component power figures are published, and the arithmetic is simple. The bad news is that once procurement is done, almost nothing can be changed: the bulk of the electricity bill is set by how many nodes you turned on, and that decision gets made once.
If you have a specification on your desk and the question of whether it will fit into your existing room, let’s work through it together. We’ll break it down by rack, by feed, and by phase, and it will immediately become clear whether you’re up against the feed, the breaker rating, or nothing at all yet.