Businesses sending their faulty RAM for repair instead of throwing it away brought two ECC RDIMM DDR5 modules, priced as much as a budget car, to a workshop called northwestrepair.
The 128 GB module was failing DDR5 training at boot, which suggested a cracked solder joint in one of the DRAM BGA packages. A reflow was enough, and the module ran flawlessly in the test system.
The 256 GB module, on the other hand, was completely dead; suspicion focused on the PMIC that powers the DRAM chips and on the SPD ROM. The SPD ROM was replaced and the old SPD backup reloaded, but nothing changed. Reflowing the suspect chips, reballing, and other steps also failed. The problem turned out to be a shorted MLCC and a faulty PMIC, likely caused by the module being dropped in the server room.
It is noted that cracked MLCCs used as filter capacitors are a common cause of failure, and that checking for short circuits speeds up diagnosis.
Why it matters
When the price of server memory reaches that of a budget car, sending a faulty module in for repair instead of throwing it away becomes a sensible option for businesses. These two cases show that while a single defect such as a cracked solder joint can be fixed with a simple reflow, repair can prove unsuccessful in multiple-fault cases combining an MLCC short circuit with a faulty PMIC; in other words, a high price does not mean every module can be saved. On the other hand, the identification of cracked MLCCs as a common cause of failure and the fact that checking for short circuits speeds up diagnosis offer workshops handling similar jobs a practical diagnostic method. The development is of direct interest to server operators using high-capacity DDR5 ECC memory and to independent repair shops; the open question is when the cost of repair on modules with multiple faults will exceed the purchase of a new module.
Background
RAM is not a new name in the FikirPilot archive: we have published 5 stories mentioning this name in the last 90 days; the most recent is dated September 29, 2026.