AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Regular Paper

Cholesky Parallel Decomposition Optimization Algorithm Based on ScaLAPACK

He Xu1,2Tao Zhou1,2Peng Li1,2( )Fang-Fang Qin3Yi-Mu Ji1,2
School of Computer Science, Nanjing University of Posts and Telecommunications, Nanjing 210023, China
Jiangsu Engineering Research Center of High Performance Computing and Intelligent Processing, Nanjing University of Posts and Telecommunications, Nanjing 210023, China
School of Science, Nanjing University of Posts and Telecommunications, Nanjing 210023, China
Show Author Information

Abstract

The scalable linear algebra package (ScaLAPACK) is a critical library for parallel computing on distributed-memory systems, enabling the development of numerous scientific applications that depend on robust linear algebra operations. However, for specific computations such as the Cholesky decomposition, the native ScaLAPACK routines are not communication-optimal and fail to fully leverage the capabilities of modern parallel architectures. This paper proposes the Parallel Cholesky Factorization (PCF) optimization algorithm designed to address these limitations within the ScaLAPACK framework. The PCF algorithm enhances performance and load balancing by strategically differentiating data partitions across processes. It involves a temporary redistribution of computational workloads to a root process, which performs a concentrated calculation before redistributing the results. This approach ensures a more balanced utilization of CPU resources. Experimental evaluation is conducted on both Intel and Kunpeng processor platforms. The first set of experiments demonstrates that the PCF algorithm achieves an average performance improvement of 30% over the native ScaLAPACK algorithm on the Intel platform under optimized process and thread configurations. The second comparative experiment shows that the PCF algorithm on the Kunpeng platform achieves an average computational efficiency increase of 200% and 50% under different thread configurations, respectively, significantly outperforming the Intel math kernel library (MKL) on the Intel platform. These results confirm that the proposed optimization effectively enhances performance and portability across diverse modern computing environments.

Electronic Supplementary Material

Download File(s)
JCST-2412-15111-Highlights.pdf (194 KB)

References

【1】
【1】
 
 
Journal of Computer Science and Technology
Pages 1009-1023

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Xu H, Zhou T, Li P, et al. Cholesky Parallel Decomposition Optimization Algorithm Based on ScaLAPACK. Journal of Computer Science and Technology, 2026, 41(3): 1009-1023. https://doi.org/10.1007/s11390-025-5111-y

9

Views

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 21 December 2024
Accepted: 23 December 2025
Published: 01 May 2026
© Institute of Computing Technology, Chinese Academy of Sciences 2026