NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization

Deng, Yanxia; Zhang, Aozhong; Gurses, Selcuk; Wang, Naigang; Yang, Zi; Yin, Penghang (August 2025, Transactions on machine learning research)

Fine-tuning large language models (LLMs) using low-rank adaptation (LoRA) has become a highly efficient approach for downstream tasks, particularly in scenarios with limited computational resources. However, applying LoRA techniques to quantized LLMs poses unique challenges due to the reduced representational precision of quantized weights. In this paper, we introduce CLoQ (Calibrated LoRA initialization for Quantized LLMs), a simplistic initialization strategy designed to overcome these challenges. Our approach focuses on minimizing the layer-wise discrepancy between the original LLM and its quantized counterpart with LoRA components during initialization. By leveraging a small calibration dataset, CLoQ quantizes a pre-trained LLM and determines the optimal LoRA components for each layer, ensuring a strong foundation for subsequent fine-tuning. A key contribution of this work is a novel theoretical result that enables the accurate and closed-form construction of these optimal LoRA components. We validate the efficacy of CLoQ across multiple tasks such as language generation, arithmetic reasoning, and commonsense reasoning, demonstrating that it consistently outperforms existing LoRA fine-tuning methods for quantized LLMs, especially at 2-bit.
more » « less
Free, publicly-accessible full text available August 17, 2026
COMQ: A Backpropagation-Free Algorithm for Post-Training Quantization

https://doi.org/10.1109/ACCESS.2025.3576737

Zhang, Aozhong; Yang, Zi; Wang, Naigang; Qi, Yingyong; Xin, Jack; Li, Xin; Yin, Penghang (June 2025, IEEE Access)

Free, publicly-accessible full text available June 6, 2026
MagR: Weight Magnitude Reduction for Enhancing Post-Training Quantization

Zhang, Aozhong; Wang, Naigang; Deng, Yanxia; Li, Xin; Yang, Zi; Yin, Penghang (December 2024, Advances in Neural Information Processing Systems 2024)

Full Text Available
Feature Affinity Assisted Knowledge Distillation and Quantization of Deep Neural Networks on Label-Free Data

https://doi.org/10.1109/ACCESS.2023.3297890

Li, Zhijian; Yang, Biao; Yin, Penghang; Qi, Yingyong; Xin, Jack (January 2023, IEEE Access)

Search for: All records