PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration

Abouelhamayed, Ahmed; Cui, Angela; Fernandez-marques, Javier; Lane, Nicholas; Abdelfattah, Mohamed

doi:10.1145/3656643

Citation Details

This content will become publicly available on March 31, 2026

PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration

Conventional multiply-accumulate (MAC) operations have long dominated computation time for deep neural networks (DNNs), especially convolutional neural networks (CNNs). Recently, product quantization (PQ) has been applied to these workloads, replacing MACs with memory lookups to pre-computed dot products. To better understand the efficiency tradeoffs of product-quantized DNNs (PQ-DNNs), we create a custom hardware accelerator to parallelize and accelerate nearest-neighbor search and dot-product lookups. Additionally, we perform an empirical study to investigate the efficiency–accuracy tradeoffs of different PQ parameterizations and training methods. We identify PQ configurations that improve performance-per-area for ResNet20 by up to 3.1×, even when compared to a highly optimized conventional DNN accelerator, with similar improvements on two additional compact DNNs. When comparing to recent PQ solutions, we outperform prior work by 4× in terms of performance-per-area with a 0.6% accuracy degradation. Finally, we reduce the bitwidth of PQ operations to investigate the impact on both hardware efficiency and accuracy. With only 2–6-bit precision on three compact DNNs, we were able to maintain DNN accuracy eliminating the need for DSPs. more »

Award ID(s):: 2303626

PAR ID:: 10580577

Author(s) / Creator(s):: Abouelhamayed, Ahmed; Cui, Angela; Fernandez-marques, Javier; Lane, Nicholas; Abdelfattah, Mohamed

Publisher / Repository:: ACM

Date Published:: 2025-03-31

Journal Name:: ACM Transactions on Reconfigurable Technology and Systems

Volume:: 18

Issue:: 1

ISSN:: 1936-7406

Page Range / eLocation ID:: 1 to 29

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
This content will become publicly available on March 31, 2026
Journal Article:
https://doi.org/10.1145/3656643

More Like this