At 1:40 a.m. on August 27, a company doing cold chain logistics in Panyu called 24 hours on duty: the only Dell PowerEdge R740 in the server room suddenly shut down, TMS dispatching system paralyzed, temperature data of more than a dozen refrigerated trucks could not be transmitted back, must be restored before 6 o 'clock tomorrow morning, otherwise a car of seafood will be scrapped.

Our engineer arrived at the scene at 2:20. The machine is R740 bought in 2021, two-way Xeon Silver 4210R, 96GB memory, two 750W platinum power supplies (model D750E-S1). Pressing the power button does not respond, the front panel diagnostic light does not light up-this is not a crash, it is completely dead.

Step 1: Don't open it in a hurry, read the iDRAC log first

R740 iDRAC is independently powered. As long as the machine is plugged in, it can be accessed even if the host is not turned on. The engineer connects the laptop directly to the dedicated network port of iDRAC. The first thing after logging in is to flip System → Overview → Logs. The result is ridiculous:

timelog contentseverity level
July 11 03:22PSU 2 input lostCritical
July 11 03:22PSU redundancy lostWarning
August 26, 23:47PSU 1 failure detectedCritical
August 27 01:38System power lostCritical

Translated into English:47 days ago, the No.2 power supply broke down. The machine has been running on the No.1 power supply for nearly one and a half months.The No. 1 power supply has been fully loaded for 24 hours since then. At the end of August, it finally failed and the whole system was directly powered off. If the alarm was replaced on July 11, the accident would not have happened.

Rack servers in data center racks
Client Room: R740 is the only server in it, without any spare parts redundancy

Step 2: Field verification to confirm the status of the two power supplies

What iDRAC says is not necessarily true. We have done physical verification:

  1. Look at the power indicator.: Unplug both power modules, LEDs 1 and 2 are amber (normally green). Amber = fault or no input.
  2. socket changing test: Plug No. 2 power supply into No. 1 socket, which was originally determined to be normal, but it still doesn't light up--Eliminate PDU socket problem and confirm No. 2 power supply body fault.
  3. Multimeter measurement No. 1 power output: 12V main output after power failure, voltage jump between 9.8V-13.5V, normal should be stable at 12V ±5%. Fan bearings also have obvious slack. Typical long-term full load aging, capacitor bulge precursor.

Conclusion: No. 2 power supply has been bad for 47 days, No. 1 power supply severe aging strike at any time. This machine was able to start up pure luck.

Step 3: Change the power supply, not just change it

750W Platinum Power Supply (D750E-S1) The market price of brand-new original parts is RMB 1300–1800, and the refurbished parts are RMB 600–800. Considering that this machine carries the scheduling business of the whole company, we suggest that customers directly replace a pair of brand-new original parts with a quote of RMB 3200 (including on-site, installation and full-load burn-in test).

There are two easy steps in the replacement process:

At 4: 50am, the TMS system came back online, 1 hour and 10 minutes before the customer's dead line. The next day we remotely reviewed the iDRAC logs and everything was clean.

Server room switch and network cable port
Switch and PDU in client rack: single-channel mains, no UPS power supply to server, this is also one of the hidden dangers

The three problems exposed by this accident, many enterprises have

Question 1: No one sees the alarm.SMTP mail alerts and SNMP traps can be configured for iDRAC, customers have never been configured. We help them configure email alerts free of charge and send them to the IT manager and boss. If your server is Dell/HP/Lenovo, this feature can be configured in 10 minutes. It is highly recommended.

Problem 2: Redundant power supplies are never tested.The meaning of dual power supply is that one can run bad, but only if "the other is good". We recommend doing a live plug test every quarter, or at least once every six months in iDRAC to see if the input voltage and fan speed of both power supplies are normal.

Question 3: Single point power supply.This R740 has two power supplies plugged into the same power PDU, and the redundancy is still zero when power failure occurs. Standard practice is to connect the two power supplies A/B respectively, and step back at least one online UPS. The 3kVA online UPS includes batteries for RMB 4000–6000, which is much cheaper than a car of seafood.

Attachment: rack server power self-test list

check itemmethodfrequency
Power LED StatusVisual, green = normal, amber = fault/no inputby the month
iDRAC/iLO/XCC LogLog in to admin to check PSU keywords in System Logsby the month
Redundant plug testTake turns to plug in two power supplies under power-on state, and confirm that there is no power lossquarterly
alarm channelConfirm that the email/SNMP alarm can be received (you can deliberately pull a power test)quarterly
Fan and air inlet dust removalCompressed air cleaning, power supply fan abnormal noise immediately replaceevery six months

If your company's server is also in the state of "no one cares when installed", you can ask us to do a free health inspection-on-site within Guangzhou, and issue a complete report containing power supply, hard drive SMART, RAID status, firmware version. No charge, the report belongs to you.