Author: b t
-
When a failed drive is not the problem
A drive marked Failed is not necessarily a failed drive. Dropped links, SMART thresholds that hide degradation, simultaneous failures, slot numbering errors and backplane faults — five common misdiagnoses and the order to work through them.
-
The hours a RAID rebuild takes
Why unrecoverable read errors make RAID 5 risky on high-capacity drives; why throughput halves while every drive reports healthy; and the two highest-risk storage operations — controller replacement and running a dual-controller array on one controller.
-
Nothing failed, the machine just got slower
A failed RAID cache battery (BBU, FBWC or supercapacitor) drops the controller from write-back to write-through and costs an order of magnitude on writes. Learn cycles, write-back that won’t re-enable after replacement, and what the battery actually protects.
-
Why the same model and capacity still won’t fit
Firmware baseline, sector format (512n/512e/4Kn), actual usable capacity, spindle speed, DIMM rank and organization, and OEM signature checks — the six dimensions of spare part interchangeability, plus the quality problems that have nothing to do with specification.
-
Memory faults, logged and unlogged
One DIMM reporting correctable errors can downclock an entire channel; wrong DIMM population order costs capacity or bandwidth silently; and the hardest case is the uncorrectable error that reboots the machine with nothing in the OS log. Diagnosing all three from the BMC.
-
SSD life is measured in writes
Three reasons an SSD slows down (SLC cache exhausted, low free space, endurance throttling); planning replacement from Percentage Used and total bytes written rather than age; and the batch firmware risk that makes a whole set fail together.
-
Packet loss, low rate, dark port
Optic vendor coding checks that keep a port dark; three reasons a 10G port negotiates at 1G; and intermittent packet loss that takes three days to trace to a patch cable. Using CRC and error counters to confirm a physical layer fault, and the cheapest-first replacement order.
-
Power and cooling: the cheapest parts protect the most expensive ones
A failed module in a redundant power supply is silent; unstable power looks exactly like failing drives; fans at full speed are usually not the fans’ fault; one stopped fan can run a CPU hot for weeks. Inspection and replacement intervals for both.
-
End of warranty, end of life, and what’s actually left
The three decisions most often made wrong in the first year out of warranty; what actually determines how many years a server has left (not its deployment date); four supply lines for discontinued parts; and the identity problem created by replacing a system board.
