GOLDYLOC: Global Optimizations & Lightweight Dynamic Logic for Concurrency

Pati, Suchita; Aga, Shaizeen; Jayasena, Nuwan; Sinclair, Matthew

doi:10.1145/3730584

Citation Details

This content will become publicly available on May 8, 2026

GOLDYLOC: Global Optimizations & Lightweight Dynamic Logic for Concurrency

Modern accelerators like GPUs increasingly execute independent operations concurrently to improve the device’s compute utilization. However, effectively harnessing it on GPUs for important primitives such as general matrix multiplications (GEMMs) remains challenging. Although modern GPUs have significant hardware and software GEMM support, their kernel implementations and optimizations typically assume each kernel executes inisolationand can utilize all GPU resources. This approach is highly efficient when kernels execute in isolation, but causes significant resource contention and slowdowns when kernels execute concurrently. Moreover, current approaches often onlystaticallyexpose and control parallelism within an application, without considering runtime information such as varying input size and concurrent applications – often exacerbating contention. These issues limit performance benefits from concurrently executing independent operations. Accordingly, we propose GOLDYLOC, which considers theglobalresources across all concurrent operations to identify performant GEMM kernels, which we call globally optimized (GO)-Kernels. GOLDYLOC also introduces a lightweight dynamic logic which considers thedynamicexecution environment for available parallelism and input sizes to execute performant combinations of concurrent GEMMs on the GPU. Overall, GOLDYLOC improves performance of concurrent GEMMs on a real GPU by up to 2 × (18% geomean per workload) versus the default concurrency approach and provides up to 2.5 × (43% geomean per workload) speedup over sequential execution. more »

Award ID(s):: 2238608

PAR ID:: 10589449

Author(s) / Creator(s):: Pati, Suchita; Aga, Shaizeen; Jayasena, Nuwan; Sinclair, Matthew

Publisher / Repository:: Association for Computing Machinery

Date Published:: 2025-05-08

Journal Name:: ACM Transactions on Architecture and Code Optimization

ISSN:: 1544-3566

Subject(s) / Keyword(s):: Concurrency-aware Execution, General Matrix Multiplication, GPGPUs

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
This content will become publicly available on May 8, 2026
Journal Article:
https://doi.org/10.1145/3730584

More Like this