First, look at the color and position of the lamp, and the fault level will come out in half.

The alarm logic of the three brands of servers (Dell PowerEdge, HP ProLiant, Lenovo ThinkSystem) is basically the same:

Key principles:Yellow lights stop immediately.. A large number of enterprise-class server components are redundant (dual power supply, RAID, ECC memory, hot plug fan), a lot of yellow lights mean "redundancy lost, repair as soon as possible", rather than "business to stop". Panic directly unplug power, but can be planned maintenance into an accident.

Two, five minutes of self-examination: remote management port is your dashboard

Don't stare at the lamp guess in the server room, log in and read the log with the external management port, and the information is much more:

brandmanagement portdefault addresssee where
DelliDRACIP on the small label pulled out of the fuselage, default user rootLog → Lifecycle Log
HPiLOBody label, default user AdministratorAHS Log/IML Integrated Management Log
associateXCC/IMMBody label, default user USERIDevent log

Do not know the IP of the management port: Dell boot press F2 to enter the BIOS to see the settings, or connect to a monitor to see the error in the POST screen; you can also press the corresponding shortcut key to enter the system settings to see. If you have not configured the IP of the management port, you will not need to run the server room next time.

III. Common alarms TOP6: from "can wait" to "stop"

① Hard drive predictive failure (PFA)-can wait, but as soon as possible

Performance: Single disk amber light flashing slowly, log "Predictive failure" or SMART alarm. The disk is still working normally in the array, but the manufacturer firmware determines that it is dying.Correct action:Take advantage of the low peak of business hot swap new disk, let array automatically rebuild. Don't wait until it really dies-RAID5 really dead one into degradation, another dead one is disaster (see our articleRAID5 double disk offline record). Key points for disk replacement: the same model and capacity are preferred, and the capacity cannot be less than the original disk; confirm that backup is available before replacement.

② Memory ECC error-frequency processing

Performance: Log "Correctable ECC"(correctable) or "Uncorrectable"(uncorrectable). Correctable error occasional once or twice: record DIMM slot position, observe, can wait for the business window to change again. Correctable error high frequency occurrence (dozens of hundreds a day) or occurrence of uncorrectable error:arranged as soon as possible, uncorrectable ECC will directly lead to downtime restart. Enterprise memory is ECC, error location to the slot (such as DIMM_A3), change the corresponding memory, pay attention to the channel pairing rules when changing.

③ Power failure-redundancy is still in operation

Performance: Dual power supply model single power amber light, log PSU fault or "Power supply AC lost". Check the simplest: power cable loose, PDU that way socket has no electricity, room that way empty jump did not jump. Really bad power supply: dual power supply model single power supply operation can support spare parts (note that the other way do not have an accident), hot plug replacement. Single power supply model according to shutdown processing, as soon as possible.

④ Fan/temperature alarm-the most undraggable type

Performance: fan crazy rotation (noise obviously increased), panel temperature light, log temperature threshold exceeded. Check first: room air conditioning is not a problem (summer Guangzhou room air conditioning failure is our high frequency visit reason), machine air inlet is not gray paste dead, fan module is not broken. When the temperature is critical, the server will reduce the frequency to protect itself, and then high on the forced shutdown-data risk.Such warnings should be dealt with immediatelyEven if it's just dust.

5 array card battery (BBU/super capacitor) aging--scheduled replacement

Performance: Log "RAID battery/capacitor degraded", write cache policy may automatically switch from Write Back to Write Through (business will slow down but safe). No emergency, just arrange spare parts replacement. Battery aging machines encounter power failure, write cache data risk loss, so do not delay too long.

④ System-level alarm light + beep--press log

The overall alarm light on the panel is usually accompanied by a beep, which is a summary prompt for any of the above problems. Some models can be remotely silenced in iDRAC/iLO. When you see the system light, check the log directly to locate the specific component. Don't be scared by the sound.

Case: An HP DL380 Gen10 "Midnight Drip"

At 1:00 a.m. on July 29, 2026, a company doing cross-border e-commerce in Panyu called us on duty: an HP DL380 Gen10 (carrying ERP and official website) in the server room began to beep, the panel yellow light, and the business was still running. In the phone, let them not touch the machine first, and report the iLO address to us for remote viewing (which was prepared during maintenance before).

The IML log of iLO shows: DIMM_B4 memory block uncorrectable ECC error once, correctable error 37, and Slot 3 hard drive SMART PFA. Two problems stacked together triggered the system alarm.

Treatment: Uncorrectable ECC once, downtime risk actually exists, cannot wait. Arrive at 2:00 a.m., spare parts are prepared in the maintenance contract (Same specification 16GB RDIMM+2TB enterprise disk). The business side switched ERP to standby solution and stopped service for 20 minutes, replaced memory module, replaced hard drive, triggered array reconstruction, and the reconstruction was suspended. The next day, the re-establishment of array was completed, and the business was not felt. This machine signed annual maintenance, and there was no additional charge for this visit and spare parts. This is the meaning of the maintenance contract:Turn midnight firefighting into planned operations with spares.

3 Tips for Business IT Administrators

  1. Set up the out-of-band management port and connect it to the monitor.iDRAC/iLO/XCC supports SNMP, email alerts, Redfish API. With email alerts, you can see problems in the morning when they occur in the middle of the night, instead of waiting for user complaints. If you don't have the energy to watch yourself, you can host them for us.IT Outsourcing & O&M7×24 monitoring alarm response.
  2. Spare parts in advance.PFA disk, memory module and power supply are the main parts of server failure. For key business machines, spare parts of the same model are kept in the server room. Failure handling has changed from "waiting for logistics in three days" to "replacing in half an hour". Annual maintenance contracts include spare parts library, which is very cost-effective.
  3. Establish alarm response ratings.Just copy our classification: Temperature/power class immediate treatment; uncorrectable ECC/double tray risk within 24 hours;PFA/battery aging scheduled treatment within a week. Write a sheet of paper next to the watch list, the person on duty to do so, do not have to call every time.

If the server lights up, take a panel photo, send the model and log screenshot to 020-39029800(same as WeChat), and help you grade first by phone. Panyu and Nansha will arrive within 2 hours.