Energy-efficient GPU SM allocation

Han, Bing-Shiun; Parekh, Kunaal; Lin, Wan-Chu; Paul, Tathagata; Gandhi, Anshul; Liu, Zhenhua

doi:10.1145/3764944.3764952

Citation Details

Energy-efficient GPU SM allocation

GPU sharing between workloads is an e!ective approach to increase GPU utilization and reduce idle power waste. To minimize resource contention under GPU sharing, current architectures allow users to allocate core GPU compute resources exclusively to workloads. However, identifying the most e''cient GPU compute resource allocation for colocated workloads is challenging, as it requires balancing potential performance degradation and power savings. This paper presents a framework for finding the most energy-e''cient compute allocation for colocated workload pairs under NVIDIA MPS using lightweight prediction models. Experimental results, using a range of training, inference, and general CUDA workloads, demonstrate that our solution outperforms the equal sharing strategy by 35%, on average, and is within 1.5% of the o#ine optimal strategy. more »