Discuss a project

Issue 08Data infrastructure

Lines and cores

A customer picks an application system and sees one line in the budget. They end up paying for four. A case study: a sizing dispute that nearly quadrupled a three-year licensing subscription, and the measurement that settled it eighteen months later.

There’s a storyline we see in projects more than any other. A customer picks an application system: anti-fraud, a reporting platform, service-channel management, the core of a new business line. The choice takes a long time and is taken seriously — comparing functionality, visiting reference sites, counting per-user and per-workstation licenses. By the time the decision is made and approved, the budget has one line in it: the cost of the system itself.

And then the day comes when that system has to actually run somewhere.

That’s when it turns out there are actually four lines: the application software, the hardware, the system licenses, and support. The customer saw the first one. They’ll pay for all four. And the most underestimated of the remaining three isn’t the hardware, as most people assume — it’s the system software licenses. Hardware, at least, is visible; you can touch it and write it off in six years. A licensing subscription isn’t visible at all, and it gets recalculated at the exact moment you decide how many cores to buy.

What follows is a case study. It’s a rare case because it came full circle: the calculation was made in January 2025, the dispute was about methodology, and eighteen months later actual metrics arrived from that very infrastructure. Sizing stories usually end with “we recommended,” and the reader is left to take it on faith. This one has a real-world check.

We’re not naming the customer or the country: Central Asia, a service with a large retail customer base. Absolute figures for customers and capacity have been replaced with ratios — the calculation logic doesn’t change as a result.

Four budget lines and when each one appearsThe decision on the application system is made on the first line. The other three follow and don't goaway.Choosing theapplicationsystemInfrastructurepurchaseYears 1–3primaryoperationYears 4–6second roundApplication software: licensesand implementationHardware: servers, storage,networkSystem licenses: virtualization,DBMS, backupHardware support and renewalsvisible when thedecision is madebecomes clear laterMain expensePartial expenseNo expenseFour budget lines and when each oneappearsThe decision on the application system is made on the firstline. The other three follow and don't go away.Choosing theapplicationsystemInfrastructurepurchaseYears 1–3primaryoperationYears 4–6second roundApplication software: licenses and implementationHardware: servers, storage, networkSystem licenses: virtualization, DBMS, backupHardware support and renewalsvisible whenthe decision ismadebecomes clear laterMain expensePartial expenseNo expense
Fig. 1. The four budget lines and when each one appears

Setting the task

Requirements that mean nothing

The infrastructure requirements section in an application system’s documentation doesn’t describe the workload — it describes the vendor’s insurance policy.

Application system documentation usually includes an infrastructure requirements section. The problem is there’s nothing to actually read there.

A typical wording looks like this: server, at least 32 cores, at least 256 GB of RAM, SSD storage. Sometimes vendors add “or equivalent” — the one honest word in the whole paragraph. No CPU generation, no clock speed, no workload profile, and no answer to the main question: is this a physical server or a virtual machine?

The difference isn’t cosmetic. Between 32 cores from 2018 and 32 cores today, the performance gap is a multiple. Between 32 physical cores and 32 vCPUs, the gap is even bigger, and it runs the other way. Buy a physical server against that line and you’ll almost certainly overpay. Size a virtual machine against it and you may come up short.

It helps to understand who writes this and why. The application vendor sets the requirements, and sets them with a margin, so that no performance scenario can be blamed on them. That’s a legitimate interest, and there’s nothing to hold against them for it. But this text can’t be used as the basis for a purchase: it doesn’t describe the workload — it describes the software vendor’s insurance policy.

The real profile is always set by application-level figures, not technical ones. For a transaction anti-fraud system, that means peak operations per second, the share of operations routed to model scoring, and how far back the training history needs to reach. For a self-service data platform, it means the number of concurrently active analysts, the size of the data marts, and whether aggregation runs on a schedule ahead of time or on the fly on request. None of these figures can be derived — or ever will be derived — from a line that reads “at least 32 cores.”

There’s also the opposite mistake, discussed less often. Requirements are understated because the vendor listed the configuration used for deployment and demos — a demo box. It genuinely runs fine for ten users and falls over at three hundred. You can spot this kind of line by its suspicious modesty: if a system meant to serve an entire bank’s front office is asking for two mid-range servers, that’s not efficiency — that’s a demo stand.

Hence the rule that opens every sizing conversation we have. Infrastructure requirements don’t come from the system’s documentation — they come from measurement. If the system is already running somewhere, we pull metrics from the live instance. If it’s new and there’s nothing to measure, we get a written workload profile from the vendor in operations, volumes, and retention periods. This isn’t something to settle verbally: a year later, there’ll be no one left to ask what anyone meant.

Oversubscription

The methodology dispute

The requirement was set in virtual resources, and the dispute was over how many physical cores should sit underneath them.

In our case, the wording was better than average. The requirement was set in virtual resources: 1,000 vCPUs per site. The application layer consisted of a container platform, relational and document databases, and supporting services. Two sites, commercial virtualization.

We calculated and proposed 5 nodes per site, each with 2 processors of 24 cores. That’s 48 cores per node and 240 physical cores per site. At a 4:1 ratio, that yields 960 vCPUs; at 5:1, 1,200. The requirement is met.

That’s where the disagreement started. The other side’s position was: there’s no magic here, a compute thread can only run on one core thread, a core has two threads, not four, so a 4:1 ratio is a marketing trick. It should be calculated one to one. A thousand vCPUs means a thousand physical cores.

The position was sincere and had its own logic. It just described the wrong system.

A one-to-one calculation doesn’t make the system any more reliable. It makes it more expensive, and it leaves part of the purchased resource permanently idle.

Oversubscription doesn’t exist because someone wants to save money on hardware — it exists because virtual machines don’t consume their allocated resources all at once. A machine with 8 vCPUs doesn’t hold 8 cores busy around the clock; it occupies them while it’s actively working and hands them back to the hypervisor scheduler in between. The entire construct of virtualization is built on exactly this. If every vCPU required a dedicated core, virtualization would deliver nothing beyond management convenience.

Worth noting separately: calculation methods that apply to bare metal, where resources really are allocated by threads, don’t carry over to virtualization. These are different allocation models, and substituting one for the other is the most common mistake in sizing disputes. And it’s an honest mistake — the person isn’t trying to cut costs or pad the number, they’re simply calculating using the scheme they know.

The recommended ratio depends on the workload type and roughly ranges from 3:1 to 6-8:1. For a mixed profile of databases and containers, a reasonable benchmark is 4-5:1. For heavy databases with strict latency requirements, it’s lowered — sometimes to 2:1, or to dedicated nodes with no oversubscription at all. A container platform that packs its own workload efficiently can comfortably run higher.

How many vCPUs per physical coreThe ratio is set by the workload type, not by caution. A one-to-one scheme is virtualization abandonedand paid for in full.typical benchmark for thisplatform6–8 : 1Test and staging environments5–8 : 1Container platform, stateless services4–6 : 1Web and API servers4–5 : 1Mixed profile: containers and databases3–5 : 1Application servers, integration bus3–4 : 1Document databases, moderate load2–4 : 1General-purpose relational databases1.5–2 : 1Databases with latency requirements1 : 1Telephony, real-time processing1 : 1The “one vCPU, one core” calculation1 : 12 : 13 : 14 : 15 : 16 : 18 : 1oversubscriptionworks confidentlylower it andcheck latencyallocate withoutoversubscriptionApproximate ranges; refine them by measuring the actual workload.How many vCPUs per physical coreThe ratio is set by the workload type, not by caution. Aone-to-one scheme is virtualization abandoned and paidfor in full.typical benchmark forthis platformTest and staging environments6–8 : 1Container platform, stateless services5–8 : 1Web and API servers4–6 : 1Mixed profile: containers and databases4–5 : 1Application servers, integration bus3–5 : 1Document databases, moderate load3–4 : 1General-purpose relational databases2–4 : 1Databases with latency requirements1.5–2 : 1Telephony, real-time processing1 : 1The “one vCPU, one core” calculation1 : 11 : 12 : 13 : 14 : 15 : 16 : 18 : 1oversubscription works confidentlylower it and check latencyallocate without oversubscriptionApproximate ranges; refine them by measuring the actualworkload.
Fig. 2. Approximate vCPU-to-physical-core oversubscription ratios by workload type

A 1:1 calculation doesn’t make the system any more reliable. It makes it more expensive, and it leaves part of the purchased resource permanently idle. And that’s not even the most unpleasant consequence.

The licensing perimeter

Cores turn into licenses

Virtualization is licensed by physical core. Everything else follows from that, including the price tag of a methodology mistake.

Now for the reason this dispute cost substantially more than it appeared to.

Commercial virtualization is licensed by physical core. Not by virtual machine, not by vCPU, not by socket, and not by server — by the processor cores physically installed in the cluster nodes. In our calculation, that’s 240 cores per site, 480 across both, and the subscription was taken with 3 years of support.

Let’s run the other scenario. Going by 1:1, you’d need 1,000 physical cores per site. At 48 cores per node, that’s roughly 21 servers per site instead of 5, and around 2,000 licenses instead of 480. Fourfold. Not a quarter more, not one and a half times more — fourfold, over a three-year subscription horizon.

The dispute was about performance calculation methodology. The cost of getting it wrong sat in a three-year licensing tail.

A counter-idea usually comes up here, and it’s worth addressing separately because it sounds convincing. The thought goes: let’s take processors with more cores, then we’ll need half as many servers, and we’ll save money.

It doesn’t work. The license follows the cores, not the servers. A 48-core processor instead of a 24-core one does roughly halve the number of nodes, but it leaves the number of licensable cores unchanged — still around a thousand per site. Meanwhile, at the time, that processor cost about 2.3 times more per unit than the one we chose. So no savings show up in any of the four lines: the hardware costs more, the license count stays the same, and support is priced off the hardware cost.

Three ways to meet the same requirementThe requirement is 1,000 vCPUs per site. The options differ not in performance but in the number ofcores you'll have to license.4–5 : 1 calculationthe accepted optionProcessor24 coresNodes per site5Cores per site240Licenses for two sites480vCPUs at the limit960 – 1,200Requirement met. Thesmallest possible licensingperimeter1 : 1 schemeon the same processorsProcessor24 coresNodes per siteabout 21Cores per siteabout 1,000Licenses for two sitesabout 2,000vCPUs at the limit1,000Four times the licenses andfour times the fleet for thesame workload1 : 1 schemeon processors twice as denseProcessor48 coresNodes per siteabout 11Cores per siteabout 1,050Licenses for two sitesabout 2,100vCPUs at the limit1,050Half the nodes, the samenumber of licenses, a pricierprocessor per unitThree ways to meet the same requirementThe requirement is 1,000 vCPUs per site. The options differnot in performance but in the number of cores you'll haveto license.4–5 : 1 calculationthe accepted optionProcessor24 coresNodes per site5Cores per site240Licenses for two sites480vCPUs at the limit960 – 1,200Requirement met. The smallest possible licensingperimeter1 : 1 schemeon the same processorsProcessor24 coresNodes per siteabout 21Cores per siteabout 1,000Licenses for two sitesabout 2,000vCPUs at the limit1,000Four times the licenses and four times the fleet for thesame workload1 : 1 schemeon processors twice as denseProcessor48 coresNodes per siteabout 11Cores per siteabout 1,050Licenses for two sitesabout 2,100vCPUs at the limit1,050Half the nodes, the same number of licenses, apricier processor per unit
Fig. 3. Three ways to meet the 1,000-vCPU-per-site requirement and their licensing consequences

This is the central takeaway of this issue. The dispute was about performance calculation methodology — a purely technical matter. But the cost of getting it wrong wasn’t in the hardware. It was in the licensing tail, which runs for three years and grows in strict proportion to the number of cores you decided to buy on day one.

And notice: the side pushing for the 1:1 scheme was defending a position that would have tripled its own budget. With the best of intentions, out of a desire to calculate more honestly. That’s usually what this looks like — not malicious intent from anyone, but a technical disagreement with financial consequences that nobody is tallying at the time.

The measurement

Tested in practice

Eighteen months later, actual metrics arrived. A dispute that had been settled in words was closed out with numbers.

Eighteen months passed. The infrastructure has been running in production, and the customer base has roughly doubled over that time. We ran an audit and pulled actual metrics from both sites.

CPU utilization at the primary site is around 45%. Virtual resources have been allocated in a volume close to the original requirement. The physical layer, meanwhile, is less than half occupied. The primary site ended up with 6 nodes instead of 5, but that doesn’t change the conclusion.

This is the answer to the question that was being argued in words back in January 2025. The 4-5:1 ratio wasn’t an optimistic assumption — in fact, it turned out to be conservative. Actual consumption landed at roughly 2.5 cores for every virtual one allocated.

Now let’s calculate the alternative. Under a 1:1 scheme, the same application workload would today be running on a fleet four times larger, with CPU utilization around 13%. Plus a licensing subscription paid three years in advance for resources that will never be used. Plus the power, rack space, and cooling for that fleet.

What the measurement showed eighteen months laterThe customer base roughly doubled over that time. The 4–5 : 1 ratio turned out to be conservative, notoptimistic.Accepted calculationactual after 18 months240 cores per site · 480 licenses1 : 1 schemethe same load todayabout 1,000 cores per site · about 2,000licenses45%13%0255075100%primary-site CPU utilization at the same application loadWhat the measurement showed eighteenmonths laterThe customer base roughly doubled over that time. The4–5 : 1 ratio turned out to be conservative, not optimistic.Accepted calculationactual after 18 months240 cores per site · 480 licenses1 : 1 schemethe same load todayabout 1,000 cores per site · about 2,000 licenses45%13%0255075100%primary-site CPU utilization at the same application load
Fig. 4. Actual CPU utilization after eighteen months versus the 1:1 calculation

The audit showed single-digit percentages at the DR site, and that’s normal: it’s on hot standby, waiting for its moment, not serving users. Its capacity is sized not by utilization but by its ability to absorb the entire primary site’s load. That’s a separate conversation, but worth mentioning — the DR site’s utilization figure regularly gets brought up as evidence that “everything is sitting idle.”

From this case, we took away a practice we now propose to everyone. At the moment sizing is agreed, we set a checkpoint: 6-12 months after go-live, pull actual metrics and compare them against the calculation. If the calculation was wrong, it will show, and there’ll be no need to look for someone to blame — there will be numbers. And the next expansion will be sized against your own figures, not someone else’s recommendations, and not an argument between two engineers.

Expansion

A new generation, and how not to pay for it twice

Two years on, the original processor is no longer available to order. Here’s what happens to the license count as a result, and how to manage it.

This same infrastructure is now expanding for further growth: the plan calls for the customer base to grow to roughly five times its original size. Nodes are being added, capacity is being added. And this is where a fork in the road appears that’s worth knowing about in advance — it shows up in any project where expansion happens a year and a half to two years after the initial delivery.

By the time expansion happens, the processor generation from the original delivery is no longer available to order. You’ll have to take the current one. And the temptation with a new generation is simple and nearly unavoidable: since we’re switching anyway, let’s take a more powerful model — it’s newer, after all.

Follow that path and here’s what happens. The core count per processor goes up — in new lineups, base models start with more cores than they did two generations ago. The licensing perimeter grows right along with it, automatically. And the application workload stays the same. You’ll pay for performance you don’t need, and keep paying for it for every year of the subscription.

We did it differently. We took the current generation with the same core count per processor — 24, as before. The license requirement comes out exactly the same as it would have on the old generation: 48 per node. Performance, meanwhile, went up, because the new generation has a higher base clock, more cache per core, and a noticeably wider memory channel. The gain came free of charge, licensing-wise.

The rule is simple and works almost every time. In a cluster running virtualization licensed by core, the processor model is chosen by core count, not by its position in the lineup. At the same core count, a new generation is almost certainly faster than the old one — and that’s the gain to take advantage of. Position in the lineup matters where there’s no per-core licensing at all: on database nodes licensed per user, on dedicated hardware, and in file and object storage systems.

Expanding onto a new generation has a second effect, less obvious than the first. The cluster becomes heterogeneous: some nodes on the old processors, some on the new. Live migration between nodes keeps working, but it usually requires enabling a compatibility mode that aligns the instruction set available to virtual machines to the lowest common denominator. In other words, the new nodes give up part of their capabilities so the cluster can stay unified.

There’s no need to worry about this: it’s the instruction set that gets aligned, not the clock speed, cache size, or memory bandwidth. The main gains of a new generation live precisely in those, and they’re fully retained. But the decision needs to be made in advance and deliberately — either a single cluster with compatibility mode, or two generation-separated clusters with independent failover planning. This should be worked out at the calculation stage, not at implementation, once the timelines are already signed.

While we’re on node configuration, a note on memory. These servers use 16 modules per server out of 32 available slots — one module per channel. You can reach the same capacity with modules half the price by filling every slot, but then memory runs at a reduced frequency. For databases, that’s noticeable, and the savings on modules come back as a slowdown, which later gets you asked to add cores. And licenses.

The 3-plus-3 horizon

The fourth line

Support doesn’t make it into the first budget because the first three years are covered by warranty. In year four, it comes back.

That leaves support. It almost never makes it into the initial budget, and the reason is understandable: the first three years are covered by warranty, and an application project’s planning horizon is rarely longer than three years.

We count the horizon not as a single term, the way presentations usually do, but as 3 years of initial operation plus 3 years of a second round. In the first three years, the hardware is new, failure risk sits with the manufacturer, and support is included in the delivery. In the second round, vendor support renewal enters the picture, and that’s separate money. The benchmark is roughly 10-15% of the cost of equivalent new hardware per year.

The three-year virtualization subscription ends at exactly this boundary, and it renews based on core count — the very count you fixed at the time of the first purchase and haven’t changed since. If you took four times more cores than you needed in year one, you’ll pay four times more on renewal too.

This leads to a non-obvious consequence: the core-count decision is the one thing in the project you make once and pay for across six years. Hardware can be partially reallocated, capacity can be added later, and you can choose not to renew support on part of the fleet. But a cluster’s licensing perimeter is tied to the hardware sitting in it, and there’s no shrinking it without decommissioning nodes.

The same logic tells you how to read proposals where licenses appear as a single line with no term attached. A line that just says “virtualization licenses,” with no subscription length or metric, means nothing: it could stand for one year or three, and any number of cores. Three things need to be asked at once — the metric, the quantity, and the term.

What can be done if the perimeter is already oversized and the money is already spent? The honest answer is: not much. You can pull some nodes out of the cluster and skip renewing their subscription, but that decision has to be made at the term boundary, not partway through, and it requires rebuilding failover. A single conversation before the purchase is far cheaper.

How the four lines behave over a 3-plus-3 horizonThe first three years are covered by warranty and a prepaid subscription. In year four everythingcomes back — at the numbers fixed in year one.Year 1Year 2Year 3Year 4Year 5Year 6Application softwareHardwareSystem licensesSupport and renewalsprimary operation, warrantysecond round: renewals andexpansionno expenselowmediumhighHow the four lines behave over a 3-plus-3horizonThe first three years are covered by warranty and aprepaid subscription. In year four everything comes back— at the numbers fixed in year one.Year 1Year 2Year 3Year 4Year 5Year 6Application softwareHardwareSystem licensesSupport and renewalsprimary operation, warrantysecond round: renewals andexpansionno expenselowmediumhigh
Fig. 5. How the four budget lines behave over a horizon of 3 years of operation plus 3 years of a second round

Checklist

What to ask before signing

Eight questions worth working through before an application system makes it into the budget.

A list of questions worth going through before the application-system line makes it into the budget. None of them requires technical background to ask.

Conclusion

In place of a conclusion

In none of these stories does anyone act in bad faith. The application vendor writes requirements with margin because they’re accountable for performance. The customer’s engineer calculates on a 1:1 basis because it seems more honest and cautious. Procurement looks at the processor’s price tag and picks the cheaper one per unit. Everyone is right in their own position.

The problem is that none of them sees all four lines at once. And the budget is built from exactly those four.

If you already have an application system picked out and the question “what will this actually cost us” on the table — let’s work it out together. Not from a line in the documentation, but from the workload profile, broken down into four lines, with a 3-plus-3 horizon. The conversation takes a couple of meetings, and the gap from the first budget usually isn’t measured in percentage points.

Discuss your project

Let’s discuss your project

Tell us about your platform or project — an engineer will reply on Telegram or by e-mail.

Message us on Telegram

Or message us on Telegram — the bot will pass your question to an engineer.