Over the past year, almost every backup project has replayed the same scene. We discuss the architecture, size the disk tier, argue about deduplication and retention depth — and by the time we reach tape, the conversation wraps up in two minutes. “Oh, that’s just the library, we’ll buy however many cartridges we need.” And that’s exactly the layer that later has to be recalculated.
There are two reasons. First, tape gets designed out of inertia, copying the configuration from a past project that had different volumes and a different media generation. Second, LTO-10 shipped in 2025, and it changed the planning arithmetic more than any generation before it: it broke backward compatibility and raised the bar for the storage network. Specifications assembled from memory now miss by more than a few percent.
This article is for the project manager and the architect who owns the solution as a whole and isn’t required to remember how many drives fit in a module. We won’t explain how a drive works. We’ll cover what determines budget and timeline: what a configuration is built from, what breaks in it, why a new generation drags along a switch replacement, and where the cartridge count actually comes from.
Its place in the architecture
Why this layer still exists
Tape does a different job in a backup system than the disk tier does, and demanding speed from it is a mistake in how the task was framed.
Every six months or so we hear that tape is dead. Usually from people who never signed off on a compliance audit report.
The conversation gets confused partly because “tape” covers two different jobs. A backup is what you restore from after a failure, and it lives for weeks. An archive is what you’re legally required to keep for years, and at best you touch it when an auditor asks. Their equipment requirements are opposite, and when these two jobs get costed as a single line item, the resulting spec is wrong in both directions at once.
Tape holds its ground on four things, and none of them is speed. Cost per terabyte at scale is a fraction of disk. A cartridge sitting in a slot draws exactly zero watts — at petabyte scale that’s a line item on the power bill, not an abstraction. Rated shelf life is thirty years or more, versus a powered-off disk measured in years. And most importantly: a cartridge pulled out of the library is physically unreachable over the network. Not “protected by policy” — unreachable. There’s air and a human being between it and an attacker.
That last point is exactly why tape lives in vaults. Everything else people call an air gap is a procedure, a setting, or an arrangement with a vendor. Those can be bypassed by compromising sufficiently high privileges. A cartridge on a shelf cannot be.
Which brings us to the main point for decision-makers: tape does a different job in a backup system than the disk tier does. Disk is responsible for getting you back up fast. Tape is responsible for making sure there’s something to restore even in the worst-case scenario. Demanding fast recovery from tape is like demanding a basement safe be convenient for grabbing lunch money.
The practical consequence: if a project’s requirements list a single recovery-time figure that’s meant to cover everything, that’s a framing error, and it needs fixing before the spec is calculated — not after.
Design and service
How the library is built and what breaks in it
All the arithmetic rests on three numbers. And the one moving part halts the entire exchange with tape.
A modern mid-range library is assembled vertically out of 3U modules, and all its arithmetic rests on three numbers.
One module gives you up to forty cartridge slots. If it’s configured with an I/O station — and it needs one, to let you insert and remove cartridges without stopping operations — thirty-five are left for data. A module holds either three half-height drives or one full-height drive plus one half-height drive. There’s no third option: a full-height drive takes up two of the three bays and can only go in the two bottom positions.
There’s one robot in the stack, and it serves every module at once. Two consequences follow from that, and they pull in opposite directions.
The good one: expanding a library is almost always cheaper than buying a second one. An expansion module is essentially a smart shelf — you don’t pay again for the robot, the controller, or the control electronics, and in the backup software the library stays one device with one shared slot pool. Two separate libraries are two separate entities in the software, separate pools, and manual job balancing.
The bad one: the robot is the single point of failure for the entire capacity at once.
And here’s a trap you won’t see in the spec sheet. The platform comes in two builds — seven modules and sixteen modules. These are different products with different part numbers, and a seven-module build cannot be converted into a sixteen-module build. The choice is made once, when you buy the base module. The advice “don’t buy a second library, expand the current one” holds exactly up to the ceiling of your specific build — and you need to check that ceiling before you promise the customer a cheap expansion.
Now, about what actually breaks. We once had the robot fail on an enterprise-class modular library — the carbon-fiber arm that rides the rails and inserts cartridges into slots with surgical precision. The library threw an error, the arm stopped responding to commands, and the entire exchange with tape came to a halt.
Replacing a part like that is not “swap a disk.” Disconnect the cabling, remove the robot from its rails, install the new one, recalibrate positioning against every wall of slots, run a full inventory, verify grip and release on cartridges. Only then back into production. One calibration mistake and the robot starts missing slots or dropping cartridges.
On an older library, on top of all that comes something that has nothing to do with technology. The platform may be discontinued, the part may be back-ordered for weeks, and an engineer with real experience on that exact model is a rarity in the market. We’ve seen cases where the procedure itself took a day, and waiting for the part took a month and a half — and for that entire month and a half, the second copy simply wasn’t being created.
That leaves two things worth building into a project deliberately, not as an afterthought: the option of a second robot on higher-end models, and a service contract with a clear reaction time for mechanical failures. And a third, minor one that gets forgotten in every other spec: cleaning cartridges. One is rated for roughly fifty cycles, and their absence tends to surface after six months of operation.
Generation change
LTO-10: what actually changed
What grew, what stayed the same, and what broke outright — item by item, without the marketing noise.
There’s a lot of marketing noise around this, so let’s go item by item — what grew, what didn’t, and what broke.
Capacity grew substantially. LTO-9 media tops out at 18 TB native. LTO-10 comes in two media types — a standard 30 TB and a premium 40 TB, both running in the same drive. That’s plus 67% or plus 122% per cartridge. All the attractive “2.2x fewer cartridges” comparisons are calculated against the premium media, and that needs to be stated explicitly, or the numbers won’t add up at procurement time.
Write speed didn’t grow. A full-height LTO-9 drive delivers 400 MB/s native; a full-height LTO-10 delivers the same 400 MB/s. The entire generational gain went into density, not speed. If someone promises a shorter backup window from moving to the new generation, ask where that comes from. Usually they’re comparing a half-height drive from the old generation (300 MB/s) against a full-height new one, meaning the gain comes from form factor, not generation.
The bus did grow, though. LTO-9 drives in this class of library connected over 8 Gb Fibre Channel. The full-height LTO-10 is rated for 32 Gb, and on compressible data it can push up to 1200 MB/s. This is the change with the most expensive consequences — the next section is about exactly that.
Cartridge capacity grew by two-thirds. Write speed — not by a single percent. The bus this speed travels over grew.
The media pre-conditioning step disappeared. On LTO-9, a fresh cartridge went through a calibration before first use that could take up to two hours and would time out jobs. LTO-10 doesn’t have it — a small thing, but it saves a working day when loading a new pool at scale.
And backward compatibility broke. Completely. An LTO-10 drive doesn’t read or write LTO-9. Through LTO-7, drives worked with the two previous generations; from LTO-8 to LTO-9, with one; LTO-10 has none left.
This isn’t a minor spec detail — it’s a reversal of procurement strategy. We’ve seen the phrase “the new generation reads old cartridges, your investment is protected” in other people’s customer justifications. It’s wrong, and the cost of that mistake is your entire previously recorded archive.
What to actually do: either keep some old-generation drives in the library for access to the accumulated archive, or migrate the data. The first means running two media pools in one device — a workable scheme, but it has to be split into logical libraries, or sooner or later the software will send a cartridge to the wrong drive. The second means reading the entire archive and rewriting it, and that’s a separate project on its own: several petabytes through a real fabric takes weeks to read out, and older media carries a higher chance of read errors.
And one detail from reviewing real specs: the “36 TB native” and “90 TB compressed” figures belong to the consortium’s roadmap, not a shipped product, and they shouldn’t make it into your calculations.
The storage network
Why a new generation drags a SAN refresh along with it
A “switch replacement” line item in a backup budget isn’t padding — it’s a direct consequence of the media generation change.
This is the section that justifies everything else. A “SAN switch replacement” line in a backup budget looks like padding — it’s actually a direct consequence of the tape generation change.
Two things need to be counted separately, and they constantly get mixed up. First — does one port have enough bandwidth for one drive. Second — are there enough ports for the whole library.
On bandwidth: an 8 Gb Fibre Channel port delivers roughly 800 MB/s. A previous-generation drive at 400 MB/s native fits into that comfortably — but on compressible data that same drive can push up to a thousand megabytes a second, and it no longer fits the port. The new generation is rated for 32 Gb and a stream of up to 1200 MB/s: on the old fabric it simply negotiates down. The customer paid for the generation and didn’t get it.
On port count: each drive occupies its own port on the switch — they don’t share a link. Twenty-four drives is twenty-four fabric ports, plus adapter ports on the media servers, plus headroom. On an old fabric with three free ports, the conversation about expanding the library ends before it starts.
A separate note on SAS. Drives also come in SAS versions, and for a small installation that’s a legitimate option: the library connects straight to the media server’s adapter, no fabric needed at all, and the whole ports conversation goes away. The limitation is elsewhere — SAS doesn’t switch: cable length is measured in meters, the number of drives is capped by the ports on the adapter, and connecting one library to two media servers already gets awkward. Once you have more than a couple of drives or more than one server, you’re back to a fabric and port arithmetic.
And a third thing people remember last: the real ceiling is usually further up the path. We once reviewed a configuration where twenty-four drives gave a rated performance of almost ten gigabytes per second — but data reached them through a couple of adapters on the media servers, and that was an entirely different number. All the arguing about “adding drives to fit the window” was pointless there: the bottleneck was neither the library nor even the fabric.
The practical sequence we recommend building into the work plan:
First, fix the drive interface. They come in Fibre Channel and SAS versions, and those are different adapters in the media server. We’ve seen a spec where the drives were ordered with FC while the server recommendations listed an external SAS adapter. Connecting a library like that is physically impossible, and it shows up at acceptance.
Next, count ports: as many ports of the required speed as there are drives, and check whether they’re free. Then total the demand and reconcile it against the bandwidth of the adapters in the media servers. Then verify there’s a free PCIe slot of the right generation in the media servers for the adapter. Only after that do you close the spec.
For a project manager the takeaway is simple: if a project involves moving to a new tape generation, the question “what about our storage network” gets asked at kickoff, not after signing. Otherwise it surfaces during implementation, once the budget is already approved.
Capacity calculation
Capacity: where the cartridge count actually comes from
There are three places where people get it wrong, and all three lean the same way: fewer cartridges get ordered than are needed.
This is where mistakes happen most, and the error always leans the same way — fewer cartridges get ordered than are needed.
First. Count native only. Tape has traditionally been sold in “compressed” terabytes: 18 TB becomes 45, 40 becomes 100. Hardware compression in the drive genuinely works — but only on data that reaches it uncompressed. A modern backup system already sends an already-deduplicated, already-compressed stream to the media. There’s nothing left to compress; the ratio will be close to one.
Any calculation that multiplies the library’s rated capacity by two and a half is an error, and it won’t surface at the project review — it’ll surface a year into operation, when cartridges run out three times earlier than planned.
Hardware compression only works on data that reached it uncompressed.
Second. Slots aren’t cartridges with data on them. We’ve seen a spec where three petabytes of archive worked out to one hundred sixty-seven cartridges, and a two-hundred-slot library got ordered “with headroom.” There’s no headroom there: those two hundred slots need to hold the archive, the pool of free cartridges for ongoing writes, cleaning cartridges, I/O-station slots, and growth over the solution’s whole lifetime.
Third, and most expensive. The number of copies is set by the regulator, not the architect. In one project, the cartridge-count calculation produced a spread: around 364 units under an optimized generational-copy scheme, versus roughly 936 under a strict reading of the “retain for 36 months” requirement. The difference — nearly six hundred cartridges — has nothing to do with the library model, the vendor, or the quality of the dedup math.
So the question “exactly which retention scheme does the regulator require, and how does your compliance function interpret it” gets asked first, before any technical debate. One unresolved point here costs more than the entire argument over whose library is better.
Where this solution ends
When tape stops being the answer
The flip side: where tape genuinely doesn’t fit, and what surfaces late and expensive when you swap it out.
A fair account requires the flip side too. Tape isn’t universal, and there are scenarios where it needs replacing.
It’s fundamentally sequential access. Restoring a single file from the middle of a cartridge means mounting, seeking, and waiting — minutes instead of seconds. Restoring a large dataset comes down to how many drives you can run in parallel.
But it’s important not to fall into the opposite mistake here. Moving off tape onto disk doesn’t by itself guarantee fast recovery. We had a case where a 1.1-terabyte virtual machine took nine hours to restore — about 34 megabytes per second. There was no tape in the picture at all: it was reading from a disk pool on ordinary 7200 RPM spinning disks.
The breakdown was more instructive than the number itself. The machine had two disks — a small system disk and a second one holding everything else. Logs showed the system disk went at around a hundred megabytes per second and finished in minutes, while the large one ran at thirty-two and ate almost the whole window. The agent restores a machine sequentially, disk by disk — but even if it could do otherwise, it wouldn’t have helped: there was nothing to parallelize, almost all the data sat on one volume. And that single stream dragged the entire chain along with it — random reads off a deduplicated pool on spinning disks, rehydration out of dedup, decompression, and the write. That’s exactly what bottomed out at thirty-something megabytes per second.
We checked and ruled out the network, the dedup database, and the hardware; the vendor confirmed the behavior through a support case.
The conclusion we repeat on every project after this: recovery time is verified by recovering, not by calculation. Backup speed tells you nothing about restore speed — they’re different chains. And if all of a machine’s data sits on one large volume, no parallelism will save it.
If tape does get replaced, it’s usually replaced with immutable object storage. And here it matters that “immutability” comes in different depths: from a lock at the media level that can’t be lifted at all, to a mechanism at the array level where the policy is changed through vendor support upon confirmation from the customer’s approved contacts — meaning the protection is already procedural, not technical. Object lock also has a soft mode, where a privileged role can remove it, and a strict mode, where no one can. Regulatory requirements call for strict, and that needs to be written into the requirements explicitly — otherwise a vendor will close the item with the softest option available.
Two things that tend to surface late and expensively on projects like this. Immutability costs capacity: for a whole set’s lock to expire at the same moment, the system keeps it self-contained, with no outside references, and dedup efficiency drops. And the compatibility matrix shapes the architecture more than the datasheet does — we’ve dealt with a case where the chosen array was certified for one backup system and wasn’t listed for a second one, and the architecture had to be split into two targets.
Checklist
Questions to ask before sizing the spec
Nine questions that get resolved in one meeting with operations and remove nearly all of the project’s risk.
- How much data actually sits on tape right now — per the backup system’s own report, not the cartridge count. And is that native or compressed: constantly confused, and the difference is several-fold.
- What generation of drives and media is installed. Not “LTO,” but the specific generation and form factor: it determines cartridge capacity, speed, and compatibility alike.
- What library configuration is in use and how many modules are already occupied. If the ceiling is reached, there’s no cheap expansion.
- How many storage-network ports are connected to the drives, and at what speed. That’s your real performance ceiling.
- What interface the drives are ordered in, and whether there’s an adapter and a free slot for it in the media server.
- What retention scheme the regulator requires and how the compliance function interprets it. The cartridge count changes by a multiple depending on the wording.
- Is there a service contract on the mechanics, and what’s its response time. The robot is the single point of failure for all the capacity at once.
- When was recovery last tested at real volume, and how long did it take. Calculated recovery time and measured recovery time are different numbers, and the difference is usually not in your favor.
- When is the media generation change planned, and is the archive migration budgeted and scheduled.
Bottom line
Instead of a conclusion
The tape layer looks like the most boring part of a backup system, and that’s exactly why it gets designed last and as an afterthought. And then it turns out to be the one place where the mistake doesn’t surface on paper — it surfaces a year later, when the cartridges have run out, the window doesn’t fit, and there’s nothing left to read the previous generation’s archive with.
With LTO-10, the cost of that mistake went up. Media capacity genuinely leapt forward, and that’s good news for the budget. But write speed stayed flat, the bus under the drive got three times wider, and the bridge to the previous generation burned down. A spec assembled from memory or copied from last year’s project isn’t just imprecise anymore — it can be non-functional.
The good news is that all nine checklist questions get resolved in one meeting with the customer’s operations team. The most expensive calculations aren’t the complicated ones — they’re the ones nobody did.