Underpowering
Translated from the Spanish original. Read in Spanish
Faults in PCs with power-supply deficiencies
Motherboard
Two places are particularly affected. Both the northbridge and the NAND gate, which acts as a pull-up, develop faults when the power supply is deficient for a long period.
The integrated pulse-width modulator circuit, a TL494 type, regulates the voltages based on the load measured at the 5-volt input. Once the power supply is overloaded, the regulator will try, unsuccessfully, to hold the voltage values; eventually the temperature in the supply rises towards the circuit’s sustainable limits. At that point the computer’s circuits have to make up for the lack of voltage by trying to draw more current, which raises the temperature of the ICs. Specifically, one of those circuits is the CPU’s VRM, which regulates the phase changes with a PWM. When the voltage drops, it has to draw more current to compensate for variations in the output waveform.
Indirectly, these faults will ultimately affect the motherboard’s northbridge. Among the typical partial faults we can find, the most notable is the one in the integrated address decoder. In general, each group of assigned addresses can fail up to 85% in the worst cases.
Another heavily affected component is the SN74-type NAND gate that acts as the pull-up for the VRM system. In the medium and long term, faults occur not because of temperature but because of the excess voltage it’s subjected to at the 5-volt input. Naturally, this is related to the type of IC: being TTL, its low tolerance to supply-voltage variations is understandable. As the datasheet shows, the maximum tolerance is +/- 0.25 volts.
In particular cases the PWM of the CPU VRM may be affected. Basically, the square wave generated to control each phase of the VRM is altered, creating transient moments in which none of the phases powers the microprocessor. Depending on how long the phase shift lasts, this will result in the microprocessor shutting down or in a serious error while processing data.
Faults in Intel microprocessors
When the microprocessor’s cooling system fails, it will reach temperatures above 55 degrees Celsius, easily hitting the limit temperature that triggers THERMTRIP. The problem lies in the changes in the silicon’s crystal structure that occur when the temperature exceeds 50 °C. Although these changes are minimal, and these microprocessors are designed to withstand maximum temperatures of 65 to 75 degrees, running the chip repeatedly for long periods close to its maximum temperatures produces significant changes in the silicon that lead to small anomalies in the nanotransistors, resulting in more instruction errors.
However, one of the most affected components is the thermal monitor’s integrated sensor. After heavy use, this sensor’s readings drift, creating a new problem. In most cases, the readings rise 5 to 10 degrees Celsius above the actual temperature. This causes yet another problem, because the processor will trigger THERMTRIP before actually reaching that temperature.
Faults in AMD microprocessors
Another kind of thermal-sensor fault occurs when sending and receiving data. In some cases the THERMTRIP signal is triggered correctly at 65 °C TJ, but the information sent to the BIOS is wrong, causing failures when it receives readings above 110 °C. Once that temperature is passed, the motherboard tries to trigger THERMTRIP and the system freezes. Since the microprocessor doesn’t cut the power — it registers no dangerous temperature rise — the motherboard cuts communication with it. The effect is similar to removing a microprocessor while the computer is running.
It’s worth noting a common pattern in the temperature reading discrepancy: the divergence grows over time.
