AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (1.3 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

Statistical reproducibility of correlation tests: Pearson, Spearman, and Kendall

Department of Mathematics, College of Science, University of Bisha, P.O. Box 551, Bisha 61922, Bisha, Saudi Arabia
Show Author Information

Abstract

Reproducibility has become a fundamental concern in modern statistical practice, yet its quantitative assessment remains limited for commonly used dependence measures. This study introduces a systematic evaluation of the reproducibility probability (RP), defined as the probability that the same statistical decision would be reached if an experiment were independently replicated under identical conditions. RP was examined for three widely used correlation tests (Pearson, Spearman, and Kendall) across different types of relationships and sample conditions. Through Monte Carlo simulations, RP was shown to provide a meaningful quantitative measure of the stability of statistical decisions across repeated experiments. Results indicated that the underlying relationship between variables, sample size, and noise level influenced reproducibility. In linear relationships, RP increased with both the strength of the true correlation and the sample size. For example, under strong linear dependence ( ρ = 0.9), RP exceeded 0.95 for n = 40 and approached 1.00 for n = 80. For weak or null correlations ( ρ = 0 or ρ = 0.3), the tests typically yielded non-significant p-values, and the corresponding RP values were generally above 0.5, reflecting stable decisions in the nonrejection area. The Pearson test demonstrated slightly higher RP in small samples due to its sensitivity to linear dependence, whereas rank-based methods achieved comparable reproducibility as the sample size increased. In contrast, under nonlinear nonmonotonic and piecewise monotonic relationships, reproducibility depended on both sample size and noise intensity. For small samples, all tests displayed highly variable RP values, while for larger samples or higher noise levels, RP values converged across methods. The results emphasized the role of RP as a reliable indicator of correlation test stability and revealed how underlying dependence patterns influenced the reproducibility of statistical results.

CLC number: 62F03, 62F05, 62G10, 62G35, 62H20

References

【1】
【1】
 
 
AIMS Mathematics
Pages 957-976

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Alshahrani ND. Statistical reproducibility of correlation tests: Pearson, Spearman, and Kendall. AIMS Mathematics, 2026, 11(1): 957-976. https://doi.org/10.3934/math.2026042

417

Views

7

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 27 October 2025
Revised: 26 December 2025
Accepted: 08 January 2026
Published: 13 January 2026
©2026 the Author(s), licensee AIMS Press.

This is an open access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0)