AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (2 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese | Open Access

Optimizing parallel matrix transpose algorithm on multi-core digital signal processors

Xiangdong PEI1Qinglin WANG1,2( )Linyu LIAO1,2Rongchun LI1,2Songzhu MEI1,2Jie LIU1,2Zhengbin PANG1
College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China
Science and Technology on Parallel and Distributed Processing Laboratory, National University of Defense Technology, Changsha 410073, China
Show Author Information

Abstract

Matrix transpose is one of the common matrix operations, which is widely employed in various fields such as signal processing, scientific computing, and deep learning. With the popularization of Phytium heterogeneous multi-core DSPs(digital signal processors) developed by National University of Defense Technology, there is a strong demand for high-performance matrix transpose implementations for Phytium multi-core DSPs. Based on the architecture of multi-core DSPs and the characteristic of matrix transpose operations, a parallel matrix transpose algorithm (called ftmMT) for matrices with different element bit widths (8 B, 4 B, and 2 B) was proposed. In ftmMT, the main optimizations include vectorization based on vector Load/Store functions, core-level parallelization based on matrix blocking, and overlapping between vectorization and memory access through implicit ping-pong methods. The experimental results show that ftmMT can significantly improve the performance of matrix transpose operations, and achieve a speedup of up to 8.99 times in comparison with the open-source transpose library HPTT running on CPU.

CLC number: TP391 Document code: A Article ID: 1001-2486(2023)01-057-10

References

【1】
【1】
 
 
Journal of National University of Defense Technology
Pages 57-66

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
PEI X, WANG Q, LIAO L, et al. Optimizing parallel matrix transpose algorithm on multi-core digital signal processors. Journal of National University of Defense Technology, 2023, 45(1): 57-66. https://doi.org/10.11887/j.cn.202301006

336

Views

2

Downloads

0

Crossref

0

Web of Science

6

Scopus

4

CSCD

Received: 09 July 2022
Published: 28 February 2023
© 2023 Journal of National University of Defense Technology

This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).