How can Windows WHEA17/AER diagnostics distinguish between PCIe Root Port, board channel, clock/power and firmware faults?
I am investigating recurring WHEA-Logger Event ID 17 PCI Express errors on Windows 11. This has progressed beyond simply identifying the device associated with Event 17.
Extensive hardware isolation and cross-OS testing has narrowed the investigation to a host/platform PCIe fault domain. However, I cannot determine whether Windows provides sufficient diagnostic information to discriminate between the remaining possible causes.
The remaining fault domains are approximately:
- CPU PCIe Root Complex / Root Port hardware — including the host-side PCIe PHY/receiver.
- Motherboard PCIe channel — traces, connectors and other physical elements between the Root Port and endpoint.
- Power delivery affecting the PCIe subsystem.
- PCIe reference clock / clocking behaviour.
- Firmware/software-controlled PCIe behaviour — UEFI/BIOS, ACPI/platform power management, link-state management, link training/retraining and related platform configuration.
The endpoint itself is deliberately not the sole focus of this question because the investigation has produced evidence on more than one PCIe path and under more than one operating system.
Windows historical WHEA records include PCIe Root Port Receiver Error, Bad TLP and Bad DLLP corrected-error conditions. A separate set of endpoint records has an Unsupported Request signature. I therefore do not believe all Event ID 17 records should automatically be treated as one phenomenon.
Linux testing has provided additional information. PCIe AER has reported Root Port Receiver Error and replay-related conditions, and receiver-margin testing has produced variable host-side measurements. However, receiver-side reporting tells me where the error was detected; it does not by itself establish whether the cause is the CPU receiver, incoming transmitter, motherboard channel, clock/power conditions or platform configuration.
This creates the diagnostic gap I am trying to solve in Windows.
What Microsoft-supported Windows diagnostics can discriminate between these five fault domains?
In particular:
- Can the complete CPER/WHEA PCI Express error section underlying Event ID 17 be retrieved and decoded using supported Microsoft tooling?
- Can I retrieve the actual PCIe AER registers/status associated with the event, including Correctable Error Status, Uncorrectable Error Status, Root Error Status and Header Log where applicable?
- Can WPR/WPA/ETW correlate a WHEA event with PCI bus configuration accesses, PnP enumeration, driver activity, ACPI activity, power-state transitions and PCIe link retraining immediately preceding the event?
- Does Windows expose enough information to distinguish an error reported by the Root Port receiver from evidence that the Root Port/CPU receiver itself is actually defective?
- Is there supported telemetry for determining the negotiated link state, speed, width, ASPM state or retraining activity at the time the WHEA event occurred?
- If Windows cannot discriminate between the physical Root Port, motherboard channel, power/clock domain and firmware/platform-state causes, what diagnostic information would Microsoft normally expect an OEM to collect to make that determination?
I am not asking for a recommendation about which component to replace. I am trying to establish the diagnostic boundary between what Windows can determine from WHEA/AER and what requires proprietary OEM or silicon-vendor instrumentation.
Ultimately, I am trying to answer a relatively simple engineering question:
Windows knows that a corrected PCIe error occurred. What information does Windows have that can tell us why it occurred, and how can that information be retrieved? I am investigating recurring WHEA-Logger Event ID 17 PCI Express errors on Windows 11. This has progressed beyond simply identifying the device associated with Event 17.
Extensive hardware isolation and cross-OS testing has narrowed the investigation to a host/platform PCIe fault domain. However, I cannot determine whether Windows provides sufficient diagnostic information to discriminate between the remaining possible causes.
The remaining fault domains are approximately:
- CPU PCIe Root Complex / Root Port hardware — including the host-side PCIe PHY/receiver.
- Motherboard PCIe channel — traces, connectors and other physical elements between the Root Port and endpoint.
- Power delivery affecting the PCIe subsystem.
- PCIe reference clock / clocking behaviour.
- Firmware/software-controlled PCIe behaviour — UEFI/BIOS, ACPI/platform power management, link-state management, link training/retraining and related platform configuration.
The endpoint itself is deliberately not the sole focus of this question because the investigation has produced evidence on more than one PCIe path and under more than one operating system.
Windows historical WHEA records include PCIe Root Port Receiver Error, Bad TLP and Bad DLLP corrected-error conditions. A separate set of endpoint records has an Unsupported Request signature. I therefore do not believe all Event ID 17 records should automatically be treated as one phenomenon.
Linux testing has provided additional information. PCIe AER has reported Root Port Receiver Error and replay-related conditions, and receiver-margin testing has produced variable host-side measurements. However, receiver-side reporting tells me where the error was detected; it does not by itself establish whether the cause is the CPU receiver, incoming transmitter, motherboard channel, clock/power conditions or platform configuration.
This creates the diagnostic gap I am trying to solve in Windows.
What Microsoft-supported Windows diagnostics can discriminate between these five fault domains?
In particular:
- Can the complete CPER/WHEA PCI Express error section underlying Event ID 17 be retrieved and decoded using supported Microsoft tooling?
- Can I retrieve the actual PCIe AER registers/status associated with the event, including Correctable Error Status, Uncorrectable Error Status, Root Error Status and Header Log where applicable?
- Can WPR/WPA/ETW correlate a WHEA event with PCI bus configuration accesses, PnP enumeration, driver activity, ACPI activity, power-state transitions and PCIe link retraining immediately preceding the event?
- Does Windows expose enough information to distinguish an error reported by the Root Port receiver from evidence that the Root Port/CPU receiver itself is actually defective?
- Is there supported telemetry for determining the negotiated link state, speed, width, ASPM state or retraining activity at the time the WHEA event occurred?
- If Windows cannot discriminate between the physical Root Port, motherboard channel, power/clock domain and firmware/platform-state causes, what diagnostic information would Microsoft normally expect an OEM to collect to make that determination?
Possible relevance of documented processor errata
Intel documents ARL068, “PCIe Gen5 Link Exit from L1 Sub-state Low Power State,” which describes a condition involving PCIe Gen5, L1 low-power state and short back-to-back PCIe configuration-space accesses/CLKREQ#.
I am not claiming that ARL068 is occurring on this system. I currently do not have the instrumentation necessary to establish the configuration-TLP timing or demonstrate the specific wake failure described by Intel.
However, because some of the investigation concerns PCIe configuration activity, power/link-state behaviour and intermittent Root Port errors, I would like to understand:
7. Does Windows contain any mitigation, detection or special handling for Intel ARL068 on affected processors?
8. Can Windows telemetry establish whether a PCIe link was in an L1 sub-state immediately before a WHEA event and whether a link exit/recovery occurred?
9. Can ETW/WPR or another Microsoft-supported facility identify PCIe configuration-space accesses sufficiently to determine whether the operating system was performing configuration accesses immediately before the event?
10. If ARL068 occurred, would Windows necessarily record Intel's documented Machine Check signature, or could additional PCIe AER/WHEA corrected errors be observed before or around such a condition?
**11. Is the mitigation/handling of processor PCIe errata such as ARL068 implemented by Windows, platform firmware/microcode, the OEM, or some combination of these?
**12. Is subscribing to Microsoft-Windows-Kernel-WHEA/Errors and decoding its CPER payload the recommended supported method for obtaining the detailed PCIe AER information behind these Event 17 records, and is there a Microsoft-supported user-mode decoder/tool for doing so?
I am not asking for a recommendation about which component to replace. I am trying to establish the diagnostic boundary between what Windows can determine from WHEA/AER and what requires proprietary OEM or silicon-vendor instrumentation.
Ultimately, I am trying to answer a relatively simple engineering question:
Windows knows that a corrected PCIe error occurred. What information does Windows have that can tell us why it occurred, and how can that information be retrieved?