A day of downtime on a GPU cluster and a day of downtime on a file server are not the same order of magnitude. Yet very few third-party providers can service AI compute hardware at all — the hard part is not swapping the part, it is judging whether it needs swapping, and whether it can be repaired instead.
We cover the full chain, from GPU board to whole cluster: spare parts and technical service.
Hardware covered
- NVIDIA HGX / SXM platforms — H100 / H200 / A100 / A800 and similar
- PCIe accelerators — L40S / A100 / RTX 4090 / 5090 and similar
- Branded AI servers — Inspur, xFusion, H3C, Lenovo, Dell, Supermicro and others
- AI fabric — InfiniBand / RoCE switches and adapters
Parts at three levels
Module level
SXM / HGX GPU modules, SXM baseboards, NVLink / NVSwitch modules.
Board level
PCIe accelerators, IB and Ethernet adapters, RAID cards, optical modules and high-speed cables, mainboards.
Chassis and ancillary
Complete AI servers, 3000W+ CRPS power modules and distribution boards, air-cooling and fan modules, memory and NVMe, chassis, backplanes, rails and looms.
The four faults that most often take a cluster down — and how we handle them

GPU falls off the bus (XID 79)
Thermal cycling after sustained full load produces cold solder joints on the power path, poor edge-connector contact, and PCIe link degradation.
Our approach: reproduce it hot to localise the fault, then repair at board level rather than replacing the whole card. The price of a GPU module and the price of a board-level repair are an order of magnitude apart.
Memory errors (ECC / row-remapping resources exhausted)
Once HBM row-remapping resources are used up the errors keep coming, and forced teardown carries very high risk.
Our approach: first establish the boundary between “repairable” and “must be replaced.” Getting that judgement wrong costs both the money and the time.
Power and cooling
3000W-class power module failure, fan module degradation, thermal paste pump-out causing sustained high temperatures.
Our approach: keep the full power and cooling parts set in stock, and check facility-side distribution and three-phase balance at the same time. A good number of “server problems” are rooted in the room’s power distribution.
Training interruptions misdiagnosed as GPU failure
Cluster training stops more often because of the fabric than the compute, and it is easily blamed on the GPUs.
Our approach: diagnose at link level, with an optical-module and high-speed-cable parts pool that supports swap-and-test on the spot.
What we can actually do
- Fault localisation: XID code analysis, fallen-off-the-bus diagnosis, ECC and row-remapping assessment
- Board-level repair and whole-unit restoration
- 48+ hours of full-load stress testing before parts leave the warehouse
- Our own 380V industrial power test environment, for full-load power-on testing of complete systems
- On-site parts caches staged in the customer’s facility against their specific models
- 24×7×4 fault response, with the part moving first
- Systems training and mentoring for the customer’s first-line engineers
In one line
Our value is not that we repair quickly — it is that when the fault happens, the part is already in stock, already tested, and ready to go straight into the machine.
Send us a GPU or AI-server fault, a part number, or a parts list to check. See Contact; a person replies within 24 hours.
