Migrations and long shutdowns

Before a data center move, everyone worries about damage in transit.

The failure statistics do not support that worry. Most equipment that fails on moving day was not broken by the move. It was the batch that was already going to fail, surfacing at the same moment.

1. The real risk is the power-down and power-up

Symptom. After a full power-down for a move or a long shutdown, some drives will not spin or are not detected on restart.

Cause. After a long stop, spindle lubricant and head parking conditions change. Restart demands more current and more torque than normal operation. A drive that is entirely healthy under continuous operation may not survive a single cold start.

What makes it worse. Drive age. A drive with five or six years on it that has never stopped may find that one cold start is the last one.

This is why “the machine was fine and it broke when we moved it.” It was not broken by the move. It had reached the point where it was running on momentum, and a cold start interrupted the momentum.

The same applies to annual maintenance outages, UPS work, and any other planned estate-wide power-down. All the same class of risk.

2. Three more things that go wrong on moving day

Disk order scrambled. Drives do not go back in the order they came out. Most controllers can rebuild an array from on-disk metadata — but not all, and not in every situation.

Cables reseated incorrectly. On multi-controller or multi-backplane systems, a wrong connection can mean an entire bank is not detected.

Array metadata lost. Especially where a controller or firmware change was folded into the move as a convenience (see section 3 of The hours a RAID rebuild takes).

What these three have in common: all are avoidable by recording beforehand, and all are expensive to recover from afterwards.

3. Preparation

Full backup. This is the one item with no substitute. The risk of a move is never zero, and the cost of a backup is known.

Record the physical layout:

  • Bay positions and disk order for every system — photograph, not just notes
  • Cable connections, labelled at both ends
  • Exported array configuration from every controller
  • Rack positions and the intended order of removal

Stage spares for critical models. Not “source it if something happens” — on site, on the day. Drives and power supply modules are the two most commonly needed.

Assess equipment condition first. Run an inspection before the move and flag anything with a deteriorating SMART trend, rising correctable error counts, or an already-failed power supply module. That is the batch most likely to surface at power-up. Replace what you can in advance.

4. Bringing it back up

Power up in stages, confirm machine by machine, not everything at once.

Two reasons:

  1. Inrush current. Starting every machine simultaneously can exceed the design headroom of the distribution. A mechanical drive’s startup current is several times its running current.
  2. Diagnosability. Bring a hundred machines up at once, three of which have a problem, and finding which three and why takes a long time.

A workable order. Network and storage first, confirmed working; then compute nodes in batches, with confirmation time between them.

What to confirm per machine. All drives detected, array status normal, no new BMC hardware events, fan and power supply state. This pass costs far less time than investigating afterwards.

5. The mixed-brand problem

Migrations and routine maintenance both run into the same issue: the equipment in one room is usually not all one brand.

The situation. International third-party maintainers are strongest on international brands, and their parts coverage and authorisation for Chinese equipment is limited. Domestic operations providers cover Chinese equipment but hold no stock overseas.

The result. One room is split across two or even three subcontractors with different response definitions — one measuring time to arrival, another time to acknowledgement. When something breaks they point at each other, and the customer coordinates between them.

Worse during a migration, which is a one-off, high-intensity, cross-brand operation. What it needs is a single point of coordination on site, not three parties each managing a subset.

Where we sit. International and domestic brands — including Huawei, Inspur, H3C, Lenovo, Sugon and Ruijie — under one contract. 80+ cities in China, 35+ countries and regions overseas, with local stock and local engineers. One contact, one set of definitions, one contract.

A migration checklist

Before (at least two weeks out)

  • Full backup, with restore verified
  • Full inspection; flag deteriorating SMART trends, rising CE counts, failed power modules
  • Replace what can be replaced in advance (drives, power supplies, cache batteries)
  • Export every array configuration
  • Photograph bay positions, disk order and cabling
  • Stage spares for critical models and count them in
  • Confirm the new room’s distribution can support staged power-up

On the day

  • Final state capture before power-down
  • Dismantle and reassemble to the record; drives and carriers stay with their chassis
  • On-site spares unpacked and accessible

Power-up

  • Staged; network and storage first, then compute
  • Confirm per machine: drives detected, array state, BMC events
  • Record everything that failed — that list is the input to the next replacement cycle

This discusses general risks in data center migrations and planned outages and does not refer to any specific project. On the relationship between parts location and achievable response times, see Five questions to ask a third-party maintainer.

← All technical notes

Discover more from 七小服 Seven Small Services

Subscribe now to keep reading and get access to the full archive.

Continue reading