AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (4.9 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese | Open Access

Quantization and pruning optimization method for attention mechanism

Yuanhong HE1,2Jingfei JIANG1,2( )Jinwei XU1,2
College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China
National Key Laboratory of Paralle and Distributed Computing, National University of Defense Technology, Changsha 410073, China
Show Author Information

Abstract

To address the significant computation and memory overhead of models based on attention mechanism, model compression techniques, such as collaborative optimization of quantization and pruning, were studied. A symmetric linear fixed point quantization method was proposed for four activation matrices of query, key, value and probability in the attention mechanism. Meanwhile, a probability matrix pruning method and a progressive pruning strategy were proposed to effectively reduce the pruning accuracy loss. Experimental results on different datasets show that for the typical attention-based model BERT, this optimization method can achieve 4 bit or 8 bit fixed point quantization and 0.93~0.98 sparsity with little or no accuracy loss, which greatly reduces the model computation and lays a strong foundation for accelerating the inference of quantized sparse models.

CLC number: TP18 Document code: A Article ID: 1001-2486(2024)01-113-08

References

【1】
【1】
 
 
Journal of National University of Defense Technology
Pages 113-120

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
HE Y, JIANG J, XU J. Quantization and pruning optimization method for attention mechanism. Journal of National University of Defense Technology, 2024, 46(1): 113-120. https://doi.org/10.11887/j.cn.202401012

582

Views

4

Downloads

0

Crossref

0

Web of Science

1

Scopus

1

CSCD

Received: 17 October 2022
Published: 28 February 2024
© 2024 Journal of National University of Defense Technology

This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).