Publications
Sort:
Open Access Research Article Issue
Statistical reproducibility of correlation tests: Pearson, Spearman, and Kendall
AIMS Mathematics 2026, 11(1): 957-976
Published: 13 January 2026
Abstract PDF (1.3 MB) Collect
Downloads:7

Reproducibility has become a fundamental concern in modern statistical practice, yet its quantitative assessment remains limited for commonly used dependence measures. This study introduces a systematic evaluation of the reproducibility probability (RP), defined as the probability that the same statistical decision would be reached if an experiment were independently replicated under identical conditions. RP was examined for three widely used correlation tests (Pearson, Spearman, and Kendall) across different types of relationships and sample conditions. Through Monte Carlo simulations, RP was shown to provide a meaningful quantitative measure of the stability of statistical decisions across repeated experiments. Results indicated that the underlying relationship between variables, sample size, and noise level influenced reproducibility. In linear relationships, RP increased with both the strength of the true correlation and the sample size. For example, under strong linear dependence ( ρ = 0.9), RP exceeded 0.95 for n = 40 and approached 1.00 for n = 80. For weak or null correlations ( ρ = 0 or ρ = 0.3), the tests typically yielded non-significant p-values, and the corresponding RP values were generally above 0.5, reflecting stable decisions in the nonrejection area. The Pearson test demonstrated slightly higher RP in small samples due to its sensitivity to linear dependence, whereas rank-based methods achieved comparable reproducibility as the sample size increased. In contrast, under nonlinear nonmonotonic and piecewise monotonic relationships, reproducibility depended on both sample size and noise intensity. For small samples, all tests displayed highly variable RP values, while for larger samples or higher noise levels, RP values converged across methods. The results emphasized the role of RP as a reliable indicator of correlation test stability and revealed how underlying dependence patterns influenced the reproducibility of statistical results.

Open Access Research Article Issue
On the reproducibility of survival quantile decisions beyond Greenwood-based precision
AIMS Mathematics 2026, 11(4): 9191-9209
Published: 03 April 2026
Abstract PDF (1,009.2 KB) Collect
Downloads:3

Greenwood-based confidence intervals are widely used to quantify uncertainty in survival quantile estimation based on the Kaplan-Meier estimator, and narrow intervals are often interpreted as evidence of stable and reliable inference. However, such numerical precision does not directly address the reproducibility of inferential conclusions under repeated sampling. The relationship between Greenwood-based confidence-interval precision and reproducibility in survival quantile inference is investigated. Reproducibility is quantified using reproducibility probability (RP), defined as the probability that a survival quantile estimate is reproduced within a specified tolerance under repeated sampling, along with its decision-based analogue RP ( D ) for two-group survival comparisons. Both measures are estimated via a nonparametric bootstrap framework under a fixed study design and sample size. Extensive simulations are conducted for single-group and two-group settings under exponential, Weibull, and lognormal survival distributions, with independent and dependent right-censoring. The results show that Greenwood-based confidence interval width is not a reliable indicator of reproducibility: Narrow intervals may coexist with low RP, whereas wider intervals may be associated with high RP, depending on the distribution, censoring mechanism, and inferential target. In two-group comparisons, decision reproducibility is driven primarily by the stability of the ordering between group-specific quantiles rather than by the numerical precision of individual estimates, and under dependent censoring, decision reproducibility can be high even when confidence intervals are wide. These findings highlight a fundamental distinction between numerical precision and inferential reproducibility in survival analysis and underscore the need to assess reproducibility alongside conventional confidence-interval reporting.

Total 2