Understanding GPU Memory Corruption at Extreme Scale: The Summit Case Study

Oles, Vladyslav; Schmedding, Anna; Ostrouchov, George; Shin, Woong; Smirni, Evgenia; Engelmann, Christian

doi:10.1145/3650200.3656615

Citation Details

Understanding GPU Memory Corruption at Extreme Scale: The Summit Case Study

GPU memory corruption and in particular double-bit errors (DBEs) remain one of the least understood aspects of HPC system reliability. Albeit rare, their occurrences always lead to job termination and can potentially cost thousands of node-hours, either from wasted com- putations or as the overhead from regular checkpointing needed to minimize the losses. As supercomputers and their components simultaneously grow in scale, density, failure rates, and environ- mental footprint, the eciency of HPC operations becomes both an imperative and a challenge. We examine DBEs using system telemetry data and logs col- lected from the Summit supercomputer, equipped with 27,648 Tesla V100 GPUs with 2nd-generation high-bandwidth memory (HBM2). Using exploratory data analysis and statistical learning, we extract several insights about memory reliability in such GPUs. We nd that GPUs with prior DBE occurrences are prone to experience them again due to otherwise harmless factors, correlate this phenomenon with GPU placement, and suggest manufacturing variability as a factor. On the general population of GPUs, we link DBEs to short- and long-term high power consumption modes while finding no signifcant correlation with higher temperatures. We also show that the workload type can be a factor in memory’s propensity to corruption. more »

Award ID(s):: 2130681 2402942

PAR ID:: 10565289

Author(s) / Creator(s):: Oles, Vladyslav; Schmedding, Anna; Ostrouchov, George; Shin, Woong; Smirni, Evgenia; Engelmann, Christian

Publisher / Repository:: ACM

Date Published:: 2024-05-30

ISBN:: 9798400706103

Page Range / eLocation ID:: 188 to 200

Subject(s) / Keyword(s):: HPC, GPU memory failures, data analysis

Format(s):: Medium: X

Location:: Kyoto Japan

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
https://doi.org/10.1145/3650200.3656615

More Like this