AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (3.5 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese | Open Access

Efficient RNN inference engine on very long vector processor

Huayou SU1,2Kangkang CHEN1,2Qianming YANG1( )
College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China
National Key Laboratory of Parallel and Distributed Computing, National University of Defense Technology, Changsha 410073, China
Show Author Information

Abstract

With the increasing depth and the inconsistent length of processing sequences, the performance optimization of RNN (recurrent neural network) on different processors makes it difficult to researchers. An efficient RNN acceleration engine was implemented for the self-developed long vector processor FT-M7032. This engine proposed a row-first matrix vector multiplication algorithm and a data-aware multi-core parallel method to improve the computational efficiency of matrix vector multiplication. It proposed a two-level kernel fusion optimization method to reduce the overhead of temporary data transmission. Optimized handwritten assembly codes for multiple operators were integrated to further tap the performance potential of long vector processors. Experiments show that the RNN engine for long-vector processors is efficient, when compared with the multi-core ARM CPU and Intel Golden CPU, the RNN-like model long short term memory networks can achieve a performance acceleration of up to 62.68 times and 3.12 times, respectively.

CLC number: TP391 Document code: A Article ID: 1001-2486(2024)01-121-10

References

【1】
【1】
 
 
Journal of National University of Defense Technology
Pages 121-130

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
SU H, CHEN K, YANG Q. Efficient RNN inference engine on very long vector processor. Journal of National University of Defense Technology, 2024, 46(1): 121-130. https://doi.org/10.11887/j.cn.202401013

297

Views

3

Downloads

0

Crossref

0

Web of Science

3

Scopus

0

CSCD

Received: 07 November 2022
Published: 28 February 2024
© 2024 Journal of National University of Defense Technology

This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).