Publications

Bmpq: bit-gradient sensitivity-driven mixed-precision quantization of dnns from scratch

Abstract

Large DNNs with mixed-precision quantization can achieve ultra-high compression while retaining high classification performance. However, because of the challenges in finding an accurate metric that can guide the optimization process, these methods either sacrifice significant performance compared to the 32-bit floating-point (FP-32) baseline or rely on a compute-expensive, iterative training policy that requires the availability of a pre-trained baseline. To address this issue, this paper presents BMPQ, a training method that uses bit gradients to analyze layer sensitivities and yield mixed-precision quantized models. BMPQ requires a single training iteration but does not need a pre-trained baseline. It uses an integer linear program (ILP) to dynamically adjust the precision of layers during training, subject to a fixed hardware budget. To evaluate the efficacy of BMPQ, we conduct extensive experiments with VGG16 …

Date
2022
Authors
Souvik Kundu, Shikai Wang, Qirui Sun, Peter A Beerel, Massoud Pedram
Conference
2022 Design, Automation & Test in Europe Conference & Exhibition (DATE)
Pages
588-591
Publisher
IEEE