SSD life is measured in writes

Mechanical drives can be replaced on a calendar. Solid-state drives cannot.

From one batch deployed on one day, some drives will be fine at three years and others will raise an endurance warning inside one. The difference is not how long they ran — it is how much was written to them. That single fact makes SSD maintenance a different discipline.

1. Slowing down is not necessarily failing

Symptom. Extremely fast when new, noticeably slower on writes after a period in service.

Three possible causes, with completely different responses:

The SLC cache is full. Many drives operate a portion of flash in a faster pseudo-SLC mode as a write cache, migrating data to the normal region afterwards. Under sustained heavy writes the cache fills and throughput falls back to native speed. This is normal behaviour, not a fault.

Free space is low. The fuller the drive, the harder garbage collection works — freeing a writable block means relocating valid data first, then erasing. Below 10–20% free, that overhead shows up clearly in latency.

Endurance is near threshold. Some models throttle deliberately as they approach their endurance limit, trading write speed for data integrity.

How to tell them apart. Two SMART values: Percentage Used and Total Bytes Written. Plenty of life left and the drive not full points at the first. Endurance near threshold points at the third.

Worth adding: read-intensive and write-intensive (or mixed-use) drives are different products with endurance ratings that differ several-fold. A read-intensive drive in a write-heavy role wears out surprisingly fast. That is a selection problem, not a quality problem.

2. Plan replacement from write volume, not age

The values that matter:

  • Percentage Used — the manufacturer’s own estimate of consumed life
  • Total Bytes Written / Host Writes — compare against the datasheet TBW or DWPD rating
  • Available Spare — remaining spare block pool; falling quickly means bad blocks accumulating

How to plan. Take the write increment over a period — a month, say — and divide the remaining writable volume by it. That estimate is far more reliable than “replace at five years”, and it gives you six months or more of warning before you need the part.

When to replace. Once the endurance warning appears, replace on a plan. Do not wait for read-only.

Why not wait. An SSD at end of endurance goes read-only. The data is intact and readable, and nothing can be written. For operations that is good news. For the service it is already an outage. On the day, the distinction does not help.

3. Skipping a firmware update can be a scheduled outage

Symptom. Certain SSD batches fail together once they reach a specific number of power-on hours.

This has happened more than once across the industry. The mechanism is broadly similar each time: a counter or state machine in firmware misbehaves under a specific condition, and the drive fails or its data becomes inaccessible. Vendors publish firmware that fixes it.

The real risk is the batch. Data center purchasing means one batch is deployed on one day, carries similar load afterwards, and accumulates power-on hours almost in lockstep. So when a trigger condition is met, they fail within a very narrow window of each other — not spread out the way random failures are.

Two or more members of one RAID set failing together is past what redundancy covers.

What to do:

  1. Track vendor firmware advisories and product bulletins
  2. Record model, firmware version, deployment date and accumulated power-on hours for the fleet
  3. When updating, do it in stages — a firmware update carries its own risk
  4. For same-batch drives, consider staggered replacement to spread their failure curves

Recording a firmware baseline matters for exactly this. Our catalogue carries one per part number, not only to match a replacement but so that when a batch-wide advisory appears, the exposure can be assessed quickly.

An inspection checklist

SSD inspection is about trends, not alerts:

  • Export Percentage Used and total bytes written for every SSD
  • Identify drives wearing noticeably faster than their peers — usually uneven load distribution
  • Anything below 20% remaining life goes into the replacement plan
  • Volumes below 20% free space: evaluate expansion or migration
  • Reconcile in-service firmware versions against current vendor advisories
  • Record deployment dates and flag same-batch groups as a correlated risk

Quarterly is frequent enough, but it has to be continuous — a single snapshot tells you nothing. The value is in the delta between two of them.

On spare parts

SSD replacement needs the full specification matched: capacity, interface (SATA, SAS, NVMe), endurance class, form factor (2.5-inch, M.2, U.2) and firmware baseline.

Seven Small Services catalogues 11,894 part numbers, 4,786 of them physically stocked. Search by brand or part number in Parts Lookup. Parts get a minimum 48-hour burn-in on receipt, a unit-level retest before shipment, and ship with a spec label and test report.


This describes common failure modes for this class of equipment, not a specific incident at a named site. Sources for the advisories and standards referenced here are listed under References.

← All technical notes

Discover more from 七小服 Seven Small Services

Subscribe now to keep reading and get access to the full archive.

Continue reading