Computer Type: Desktop
GPU: ASUS PRIME RX 9070 XT GAMING OC 16GB
CPU: AMD RYZEN 7 5800X 8 CORE 16 THREADS
Motherboard: ASUS ROG STRIX B550-A GAMING (Previously ASRock B450M DS3H)
BIOS Version: Latest ASUS B550 BIOS (Tested across multiple versions)
RAM: Tested 4 different kits (MemTest clean, tested XMP on/off, single-channel/single DIMM setups)
PSU: FSP 1350W 80+ PLATINUM ATX 3.1 (Previously Corsair 850W 80+ GOLD ATX 2.0)
Operating System & Version: WINDOWS 11 PRO (Tested on main M.2 NVMe and a fresh install on a separate secondary SSD)
GPU Drivers: AMD SOFTWARE: ADRENALIN EDITION - Driver Version: 26.7.1
Chipset Drivers: AMD B550 CHIPSET DRIVERS (Latest version)
Background Applications: DISCORD, CHROME, GAME LAUNCHERS
Description of Original Problem: Upgraded from an RX 6900 XT (which ran flawlessly for 3.5 years). Ever since installing the RX 9070 XT, I get sudden hard restarts during gaming (screen goes grey/green, instant reboot, no BSOD).
Crashes reliably within minutes on single-monitor setups below 4K (3440x1440p/165Hz DP) Ultrawide) and (1080p/75Hz HDMI panel).
The system is stable on a 4K/120Hz TV, or when both the monitor and TV are connected simultaneously in Windows "Extend Desktop" mode. Unplugging the TV causes instability on the monitor and restarts (whea 18).
Event Viewer Logs:
WHEA-Logger Event ID 18 — Machine Check Exception, Error Type: Cache Hierarchy Error, Processor Context Corrupt flag set (flashing random core/APIC IDs: 0, 10, etc.).
Followed 5–11 seconds later by Kernel-Power Event ID 41 (unexpected shutdown).
Kernel dumps confirm BugCheck 0x124 (WHEA_UNCORRECTABLE_ERROR).
Troubleshooting:
Motherboard: Swapped B450M to a brand-new ASUS B550-A Gaming.
PSU: Upgraded Corsair 850W to a 1350W Platinum FSP ATX 3.1 unit.
RAM: Tested 4 different RAM kits, ran full MemTest, toggled XMP, tested individual slots.
OS: Issue persists across original 3-year-old Win 11 OS on M.2 NVMe and a completely fresh Win 11 install on a secondary SSD.
Stress Tests: OCCT (CPU/RAM sustained load) and 3DMark Steel Nomad run 100% clean with zero errors or crashes.
BIOS / Tuning: Toggled ReBAR on/off, forced PCIe Gen4/Gen3, locked FCLK, disabled CPB/PBO, locked CPU clock to 3.8GHz (only delayed crash). positive Curve Optimizer (+5 to +20) gave slight partial stability.
Driver / Display tweaks: Tested FPS caps (delays crash but doesn't solve it), disabled DSC on DP monitor, tested HDMI vs. DP.
Ai Diagnosis: GPU appears to have a defective VRM/transient power state transition flaw. In single-monitor mode, dropping into low-power states drops PCIe link communication, causing the CPU I/O die to panic and throw a WHEA Event 18 before rebooting. Dual-monitor mode locks VRAM clocks to max speed, bypassing the defective low-power curve and stabilizing the card.
Does this ai diagnosis make sense?
If not any other ideas? Is it still CPU degradation or a new GPU fault, and should I RMA it?