Discuss a project

Issue 01Data infrastructure

Hardware for BigData

Where the array belongs, where local disks belong, and where virtualization shouldn't happen at all. A breakdown of three forks that shape the architecture and budget of a data platform.

The past six months have brought a wave of BigData and Data Lake projects. We won’t touch the software side — everyone picks their own stack there. But on hardware, phasing, and sizing we keep seeing the same set of rakes people step on with impressive regularity. So we decided to break it down separately.

It usually starts with a specific request. A customer comes in: “we need 150 terabytes of storage for an analytics platform.” Where does 150 come from? We open the farm diagram, and there it is — someone added up the disks of every VM, and that’s how the round number appeared.

On its own, that number means nothing. Behind it hide three questions, and until they’re answered, any specification is guesswork. Expensive guesswork, too: the mistake doesn’t surface at procurement — it surfaces six months later, when the space has run out and the budget is already spent.

Nature of the data

Question one: what actually compresses

The array works with whatever reaches it. Anything compressed or encrypted inside the virtual machine arrives as an already-finished stream.

Let’s start from the end, because this is the most expensive mistake of the three.

The logic is trivially simple. The array works with whatever reaches it. If data is compressed or encrypted inside the virtual machine, what reaches the array is an already high-entropy stream — there’s nothing left to compress. Neither compression nor compaction will help, because the work has already been done a layer up.

Let’s break it down further, because there are a lot of nuances here.

What compresses on the array and what doesn'tThe storage system works with whatever reaches it. Anything compressed or encrypted inside the VMarrives as a finished stream.3–10 : 1Text logs, configurations, source code3–8 : 1Airflow, Superset, Metabase metadata DBs3–8 : 1CSV, JSON, XML, uncompressed DBMSdumps3–5 : 1Kafka with compression.type = none2.5–5 : 1NiFi content repository (text stream)2–4 : 1PostgreSQL, MySQL without tablecompression2–4 : 1OS images and VM system disks2–4 : 1HDFS with uncompressed text files2–3.5 : 1Avro without a codec, JSON Lines1.3–1.8 : 1Elasticsearch, OpenSearch indexes1.5–2.5 : 1Parquet / ORC without a codec1.05–1.2 : 1DBMS with internal page compression1.05–1.15 : 1Parquet / ORC + snappy, zstd, gzip1.05–1.15 : 1ClickHouse MergeTree (LZ4 / ZSTD)1.05–1.15 : 1Kafka with compression.type = lz4 / zstd1–1.1 : 1Deduplicated backups1–1.05 : 1Archives: zip, gz, 7z, zst1–1.05 : 1Media: H.264/H.265, jpeg, mp3, flac1 : 1Encrypted inside the VM or by theapplication1 : 12 : 14 : 16 : 18 : 110 : 1the array gives areal gaindepends on thecodecdata arrivesalready packedApproximate ranges; verify against your actual data.What compresses on the array and whatdoesn'tThe storage system works with whatever reaches it.Anything compressed or encrypted inside the VM arrivesas a finished stream.Text logs, configurations, source code3–10 : 1Airflow, Superset, Metabase metadata DBs3–8 : 1CSV, JSON, XML, uncompressed DBMS dumps3–8 : 1Kafka with compression.type = none3–5 : 1NiFi content repository (text stream)2.5–5 : 1PostgreSQL, MySQL without table compression2–4 : 1OS images and VM system disks2–4 : 1HDFS with uncompressed text files2–4 : 1Avro without a codec, JSON Lines2–3.5 : 1Elasticsearch, OpenSearch indexes1.3–1.8 : 1Parquet / ORC without a codec1.5–2.5 : 1DBMS with internal page compression1.05–1.2 : 1Parquet / ORC + snappy, zstd, gzip1.05–1.15 : 1ClickHouse MergeTree (LZ4 / ZSTD)1.05–1.15 : 1Kafka with compression.type = lz4 / zstd1.05–1.15 : 1Deduplicated backups1–1.1 : 1Archives: zip, gz, 7z, zst1–1.05 : 1Media: H.264/H.265, jpeg, mp3, flac1–1.05 : 1Encrypted inside the VM or by the application1 : 11 : 14 : 16 : 18 : 110 : 1the array gives a real gaindepends on the codecdata arrives already packedApproximate ranges; verify against your actual data.
Fig. 1. Approximate data-reduction ratios by data type

What will never compress. Anything encrypted at the application or OS level — TDE in the DBMS, LUKS and BitLocker inside the VM, SSE in object storage. Exactly 1:1, no exceptions. A good cipher produces output that’s statistically indistinguishable from random.

An important caveat that gets confused constantly: SED drives and volume encryption done by the storage array itself have no effect on the ratio at all. They operate after compression. The problem is only encryption above the array level.

Next come already-compressed formats. Archives, media, deduplicated backups. Here it’s 1.0–1.05:1, and you simply can’t count on them when sizing.

The weighted ratio for such a platform comes out at 1.1–1.2 to 1. Yet the capacity calculation still assumes 2 to 1.

What won’t compress even though it looks innocent. This is where the real surprises live.

ClickHouse. MergeTree compresses data with its own codecs — LZ4 by default, often ZSTD. What lands on the array is already finished. Someone plugs in 2:1 “because it’s a database” — and misses by a factor of two.

Kafka. Depends on the producers’ compression.type. With snappy, lz4, or zstd, it’s already compressed on the client, and the array gets 1.05–1.15:1. With none, log segments compress beautifully, 3–5:1. This is the one heavy component where the answer might actually work in your favor. Go and ask.

Parquet and ORC with snappy or zstd. Data lake classics — and classically incompressible. If that’s what sits behind the term “object storage” in your case, forget deduplication as a sizing factor.

Databases with internal page compression — Oracle Advanced Compression, HCC, SQL Server page compression. A separate fork: sometimes it’s more cost-effective to turn compression off in the DBMS and hand it to the array, taking the load off server CPUs. But that’s a decision requiring a mandatory performance test, not a “let’s just try it in production.”

What compresses beautifully. Text logs, Airflow and Superset metadata databases, configuration files, source code — 3–10:1. Uncompressed CSV, JSON, XML, dumps. PostgreSQL without table compression. OS images of identical VMs.

Deduplication is a separate mechanism, and people forget about it. It works independently of compressibility. Two hundred identical VMs cloned from one template: every block is incompressible, but blocks repeat, and dedup delivers a great result. Conversely, a unique, incompressible dataset won’t yield to compression or dedup either. There it’s an honest 1:1, and no vendor guarantee will change that.

Now the arithmetic this was all for. Take a typical stack: ClickHouse nodes, object storage, Kafka brokers. In a normal project that’s around 90% of the total allocated volume. And all three compress data themselves. The nicely compressible part — metadata databases, logs, configs — is a few percentage points, and it has no effect on the overall ratio.

The weighted ratio for such a platform comes out at 1.1–1.2:1. Yet the capacity calculation still uses 2:1, because that’s what’s written in the standard virtualization sizer. The gap between these two numbers is exactly the delta someone will discover six months later. Usually not the person who signed off on the spec.

Placement architecture

Question two: where the data should physically live

Engines with their own replication belong on local disks. The array is needed where fast failover matters.

ClickHouse, Kafka, Greenplum, HDFS, and Elasticsearch were designed for local disks. They have their own replication, their own view of data placement, their own node-failure recovery mechanisms. This is shared-nothing architecture, and it’s not about the array.

Put something like that on shared storage and you get double protection: RAID on the array plus the engine’s own replicas. Capacity gets spent twice. And the bottleneck moves into the fabric, where every node in the cluster starts competing for the same ports. Instead of a distributed system, you end up with a distributed system that has one shared point of failure and one shared queue.

The reverse gets forgotten more often, and that’s a mistake too. Master nodes, catalogs, schema registries, metadata databases, VM images — they don’t need throughput. They need fast failover when a host dies, snapshots, and predictable recovery. That’s exactly what the array is bought for. Keeping them on local disks means voluntarily giving up HA in the one place where it comes almost for free.

A separate note on the object tier, which in our part of the world somehow always gets considered last. ClickHouse, Kafka, Greenplum, and Spark can all offload cold data to S3 today. ClickHouse — via S3 disk and storage policies. Kafka — via tiered storage. This is usually the cheapest way to handle volume growth, and it’s almost always left out of the initial discussion. Then it turns out 60% of the data hasn’t been touched in six months and could easily have been sitting on QLC capacity at a third of the price.

Where the analytics platform's data physically livesEngines with their own replication live on local disks. The array is needed where fast failover matters.Analytics engine dataKafka — log segmentsClickHouse — data partsGreenplum — segment dataHDFS DataNodeSpark — shuffle and temporaryfilesTrino / Presto — spillElasticsearch / OpenSearch —shardsPipeline glueAirflow — metadata DB, task logsNiFi — flowfile and contentrepositoryKafka Connect, DebeziumApache Karaf — bundles, deploy,logsSchema RegistrySuperset, Metabase — metadataDBsZooKeeper / etcdPlatform and shared dataMaster nodes and catalogsOS images and VM system disksShared datasets for model trainingCold tier, archiveBackupsLocal NVMein the serverBlock storage(FC / iSCSI)NAS(NFS / SMB)Object storage(S3)RecommendedAcceptable, with caveatsNot recommendedWhere the analytics platform's dataphysically livesEngines with their own replication live on local disks. Thearray is needed where fast failover matters.Analytics engine dataKafka — log segmentsClickHouse — data partsGreenplum — segment dataHDFS DataNodeSpark — shuffle and temporary filesTrino / Presto — spillElasticsearch / OpenSearch — shardsPipeline glueAirflow — metadata DB, task logsNiFi — flowfile and content repositoryKafka Connect, DebeziumApache Karaf — bundles, deploy, logsSchema RegistrySuperset, Metabase — metadata DBsZooKeeper / etcdPlatform and shared dataMaster nodes and catalogsOS images and VM system disksShared datasets for model trainingCold tier, archiveBackupsLocal NVMein the serverBlock storage(FC / iSCSI)NAS(NFS / SMB)Objectstorage(S3)RecommendedAcceptable, with caveatsNot recommended
Fig. 2. Placement matrix: local disks, block storage, NAS, and object storage

Virtualization

Question three: what to virtualize

The line isn’t drawn by product name, but by required throughput and acceptable latency jitter.

There’s rarely an honest yes-or-no answer here. The line isn’t drawn by product name, but by required throughput and by how critical latency jitter is.

What makes sense to virtualizeThe line isn't drawn by product name, but by required throughput and acceptable latency jitter.Virtualizewithout caveatsAirflow, Superset,Metabase metadata DBsApache Karaf, OSGicontainersKafka Connect,DebeziumSchema Registry,metadata registriesMonitoring, orchestration,CIMaster nodes andcatalogsDev and testenvironmentsLow I/O load; the value is inlive migration, snapshotsand fast recoveryVirtualizeif conditions are metClickHouse undermoderate loadKafka with a modeststreamNiFi — small pipelinesTrino / Presto coordinatorSpark driver, smallexecutorsZooKeeper, etcdS3 gateways and proxiesvCPU reservation withoutovercommit, NUMA pinning,a separate datastore,monitoring of ready timeand co-stopKeepon physical serversClickHouse under heavyqueriesKafka with a high writestreamGreenplum segmentsHDFS DataNodeSpark executors on largevolumesElasticsearch data nodesGPU model-trainingnodesThe hypervisor takes awaylatency predictability, whilethe engine counts on directaccess to NVMeWhat makes sense to virtualizeThe line isn't drawn by product name, but by requiredthroughput and acceptable latency jitter.Virtualizewithout caveatsAirflow, Superset, Metabase metadata DBsApache Karaf, OSGi containersKafka Connect, DebeziumSchema Registry, metadata registriesMonitoring, orchestration, CIMaster nodes and catalogsDev and test environmentsLow I/O load; the value is in live migration, snapshotsand fast recoveryVirtualizeif conditions are metClickHouse under moderate loadKafka with a modest streamNiFi — small pipelinesTrino / Presto coordinatorSpark driver, small executorsZooKeeper, etcdS3 gateways and proxiesvCPU reservation without overcommit, NUMApinning, a separate datastore, monitoring of readytime and co-stopKeepon physical serversClickHouse under heavy queriesKafka with a high write streamGreenplum segmentsHDFS DataNodeSpark executors on large volumesElasticsearch data nodesGPU model-training nodesThe hypervisor takes away latency predictability,while the engine counts on direct access to NVMe
Fig. 3. Three zones of virtualization suitability by component

Metadata databases, connectors, orchestration, registries, Karaf with its bundles, monitoring — these virtualize without a second thought. I/O load is low, and the payoff from live migration, snapshots, and fast recovery outweighs everything else. There’s nothing to argue about here.

Greenplum segments, DataNode, Kafka under a heavy write load, Elasticsearch data nodes — these usually stay on bare metal. Not because they “won’t work” — they’ll work fine. But the engine expects direct access to NVMe, and the hypervisor hands back predictability with caveats. You won’t see the difference in tests; you will see it in production under load.

Monitoring shows everything is fine, and the system is unusable.

Between these two extremes lies the largest and most interesting zone, where configuration decides everything. And here’s an observation from practice: problems in the middle zone almost never come from a shortage of resources. They come from vCPU overcommit and from the VM being smeared across sockets.

The symptom is painfully familiar: scheduler wait time climbs, CPU utilization looks modest, every graph is green, and users complain about stutter. Monitoring shows everything is fine, and the system is unusable. Ready time and co-stop are the first thing to check, not the last.

The second classic source is a shared datastore. An analytics VM eats up the queue while something latency-sensitive lives right next to it on the same LUN. Separate them at the design stage, not once it starts hurting.

Data pipelines

The glue: the last thing anyone remembers

Averaged, ward-wide sizing breaks exactly on the glue components: every one of them has its own requirements profile.

Everyone remembers ClickHouse and Kafka. Pipelines — almost never, and that’s a dozen components, each with its own character.

Load profiles: what each component asks forOne and the same cluster serves very different requirements. Averaged, ward-wide sizing breaksexactly here.CPURAMDisk:IOPSDisk:throughputNetworkLatencysensitivityVolumegrowthKafka brokerClickHouseGreenplum segmentSpark executorTrino / Presto workerElasticsearch data nodeHDFS DataNodeNiFiApache Karaf / OSGiKafka Connect, DebeziumAirflow scheduler + workerSuperset, MetabaseZooKeeper / etcdObject tier (S3)lowmediumhighLoad profiles: what each component asksforOne and the same cluster serves very differentrequirements. Averaged, ward-wide sizing breaks exactlyhere.CPURAMDisk: IOPSDisk: throughputNetworkLatency sensitivityVolume growthKafka brokerClickHouseGreenplum segmentSpark executorTrino / Presto workerElasticsearch data nodeHDFS DataNodeNiFiApache Karaf / OSGiKafka Connect, DebeziumAirflow scheduler + workerSuperset, MetabaseZooKeeper / etcdObject tier (S3)lowmediumhigh
Fig. 4. Requirement profiles of platform and pipeline-glue components

Kafka Connect and Debezium. A CDC stream out of transactional databases. The disk needs almost nothing, but memory for buffers and a stable network are essential. Latency suffers badly if this ends up on an overcommitted host.

NiFi. This is the one that’s systematically underestimated. It has three repositories — flowfile, content, provenance — and all three write actively. Under load, the content repository reliably generates more IOPS than people expect. Provenance grows quietly and one day fills the disk completely. Local NVMe, separate volumes per repository, fill-level monitoring — none of this is a luxury.

Apache Karaf and OSGi containers in general. Modest in resource needs, lives happily on a VM. What it actually needs is a fast restart and a snapshot before deploying bundles, because hot rollbacks happen more often than anyone would like. This is an argument for virtualization, not against it.

Airflow. Metadata database, task logs, DAGs. Light on resources, but logs grow linearly and endlessly if rotation isn’t configured. We’ve seen a 200-gigabyte Airflow metadata database bring down the scheduler simply because nobody cleaned up the run history.

Schema Registry, ZooKeeper, etcd. The volumes are trivial, but ZooKeeper and etcd are extremely sensitive to fsync latency. A slow disk under etcd, and the cluster starts re-electing its leader for no apparent reason. You can’t just put them “wherever there’s room.”

Trino and Presto. The coordinator virtualizes without issue. Workers want memory and network, and for spill — a local fast disk. If spill goes out over a network volume, heavy joins turn into a pumpkin.

Superset and Metabase. The BI layer, small metadata databases, but query caches grow. They go on a VM, live on the array, and raise no questions.

The point of this whole list is simple: averaged, ward-wide sizing breaks exactly on the glue components. You can’t take the farm’s total volume and divide it by the number of servers. One component needs memory, another needs IOPS, a third needs low fsync latency at a tiny volume. You have to size by profile.

Hardware map

Now the hardware: what fits what

A platform is almost always assembled from several types of hardware. Trying to get by with one is the source of most problems.

Let’s move on to specifics, because abstractions are fine right up until you sign the specification.

Let’s say this up front: a platform is almost always assembled from several types of hardware. Trying to get by with one type is exactly the source of most of the problems we’ve described above.

Hardware map: which layer runs on which hardwareA platform is almost always assembled from three types of hardware. Trying to get by with one is thesource of most problems.ClickHouse, Kafka,Greenplum dataHDFS DataNode, SparkexecutorElasticsearch data nodesGPU model-training nodesGlue: Airflow, NiFi, Karaf, CDCMaster nodes, catalogs,registriesMetadata DBs andsupporting DBMSsBI layer: Superset, MetabaseDatastores for virtualizationShared datasets, NFS fortrainingS3 object tier for dataCold tier and archiveBackupsThinkAgile HX(Nutanix HCI)ThinkSystembare metalDG Series(QLC,capacity)DM Series(NVMe,speed)DE / DS Series(block SAN)Main optionPossible, with caveatsNot suitableHardware map: which layer runs on whichhardwareA platform is almost always assembled from three types ofhardware. Trying to get by with one is the source of mostproblems.ClickHouse, Kafka, Greenplum dataHDFS DataNode, Spark executorElasticsearch data nodesGPU model-training nodesGlue: Airflow, NiFi, Karaf, CDCMaster nodes, catalogs, registriesMetadata DBs and supporting DBMSsBI layer: Superset, MetabaseDatastores for virtualizationShared datasets, NFS for trainingS3 object tier for dataCold tier and archiveBackupsThinkAgileHX(NutanixHCI)ThinkSystembare metalDG Series(QLC,capacity)DM Series(NVMe,speed)DE / DSSeries(blockSAN)Main optionPossible, with caveatsNot suitable
Fig. 5. Matching platform layers to hardware types

Lenovo ThinkAgile HX (Nutanix) — for the glue and the platform layer.

This is where HCI plays to its strengths. Metadata databases, Airflow, NiFi for small pipelines, Karaf, CDC connectors, registries, the BI layer, master nodes, dev and test environments. Everything where the value lies in manageability, snapshots, and fast recovery, not in squeezing out the last microsecond.

Nutanix Unified Storage deserves a separate look. It has an S3-compatible object layer that covers the cold tier for analytics right on the same cluster, with no separate hardware. Plus Files for shared datasets and Volumes for block access. For mid-sized platforms this often removes the need for a separate storage array entirely.

Currently: the ThinkAgile HX650 V4 on Intel Xeon 6 is the workhorse for this layer. The HX650a V4 — if that same cluster also needs GPUs for inference.

Lenovo ThinkSystem bare metal — for engine data.

Everything that’s shared-nothing and disk-hungry: ClickHouse nodes, Kafka brokers under load, Greenplum segments, DataNode, Spark and Trino workers, Elasticsearch data nodes, model-training nodes.

The SR650 V4 in 2U delivers up to 36 NVMe drives and up to 10 PCIe Gen5 slots — plenty for an analytics node. Xeon 6 with P-cores, up to 86 cores per socket, MRDIMM and CXL support. For ClickHouse and Greenplum, where everything comes down to memory and its bandwidth, that matters noticeably.

An important point that often gets missed: analytics doesn’t need the fastest processor, it needs the right balance of cores, memory, and NVMe lanes. A top-tier CPU starved of disk lanes will just sit waiting for data. This is exactly the case where skimping on disks makes overpaying for the processor pointless.

Lenovo DG Series (QLC, capacity) — for the cold tier, archive, and backups.

QLC flash on ONTAP. Cheap capacity that’s still flash. The ideal spot for the cold tier, shared datasets over NFS, and backup target storage. It has S3 right on the array, which removes the need for separate object storage for cold data.

One nuance from practice: usable capacity on these arrays runs roughly a third below raw — right-sizing, ADP partitioning, RAID-DP dual parity, WAFL reserve. That’s normal for enterprise-class all-flash, but it needs to be built into the calculation from the start, not discovered at commissioning. And you need to count in the same units as the requirement: sizers calculate in base-2 but the spec gets signed in base-10 — roughly a 10% gap.

Lenovo DM Series (NVMe, performance) — for whatever needs speed and HA.

Master nodes, catalogs, metadata databases, datastores for virtualization, shared datasets where latency matters. Unified — meaning both block and file. When you need fast NFS for shared training data, this is where it goes.

Lenovo DE and DS Series (block SAN) — for datastores and supporting DBMSs.

Classic block access. DE is midrange, available hybrid and all-flash; DS is block all-flash. Datastores for virtualization, supporting databases — wherever you don’t need file access or ONTAP features.

What not to do. Don’t put Greenplum segments and ClickHouse nodes on block storage, no matter how fast it is. Don’t keep metadata databases and master nodes on local disks with no HA. Don’t buy a single array “for everything” and then have to explain why analytics is slow and backup won’t fit in its window.

Procurement phasing

On phasing, since we’re here

A data platform is built iteratively, yet procurement somehow keeps getting planned as one big delivery three years out.

A separate headache — not about hardware, but about the same money.

A data platform is built iteratively. First a test bench and a pilot on a couple of nodes, then production, then expansion for the real workload. The problem is that procurement somehow keeps getting planned as one big delivery “three years out.”

The result is predictable: half the capacity sits idle heating the room in year one, and by year three it turns out the workload profile changed and something completely different was actually needed. On top of that, the hardware ages while the warranty clock runs from delivery, not from go-live.

It’s smarter to build in expandability and buy for the current phase plus a clear horizon. Fortunately, HX, ThinkSystem, and DG all expand cleanly with nodes and disks, with no service interruption.

That said, there’s a counterargument that’s louder than usual in 2026: AI-driven demand growth has made lead times unpredictable. Long deliberation has gotten expensive. So the balance between “don’t overpay now” and “don’t wait six months later” is something everyone has to find for themselves. But it has to be a conscious search, not the principle of “bought extra just in case.”

Checklist

Questions to ask before sizing the spec

Six questions that turn a round number into a calculation you can defend.

A short list that saves weeks of back-and-forth:

  1. Codecs. What’s configured in ClickHouse MergeTree, what compression.type the Kafka producers use, what parquet in the object tier is compressed with.
  2. Encryption. Is there TDE, LUKS, SSE — anything above the array level.
  3. Actual fill level. How much is written on the disks today, not how much is allocated. The gap is usually a multiple.
  4. Growth forecast. Over the planning horizon, broken down by component. Kafka with a 7-day retention and ClickHouse holding three years of history grow in completely different ways.
  5. Access profile. What share of the data is actually read in a month. This answers the question of whether a cold tier is needed.
  6. Recovery requirements. RPO and RTO for each layer. The Airflow metadata database and raw Kafka data have different ones, and protecting them identically is overspending.

Checking compressibility estimates isn’t hard: most vendors have utilities that scan a directory of real data and produce a forecast before you even migrate anything. Half a day of work against months of explaining why the space ran out.

Bottom line

Instead of a conclusion

A round number in a request isn’t a requirement — it’s a symptom. It means someone added up VM disks and didn’t ask a single one of the six questions above.

The good news: all six can be settled in a couple of meetings with the platform team. The bad news: without them, any specification is a bet, not a calculation. And it gets tested not on paper but in production, six months later, when changing anything is already expensive.

Planning a data platform? Get in touch — we’ll work through your specific configuration. And while we’re at it, we’ll calculate what will actually compress for you.

Discuss your project

Let’s discuss your project

Tell us about your platform or project — an engineer will reply on Telegram or by e-mail.

Message us on Telegram

Or message us on Telegram — the bot will pass your question to an engineer.