AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (1.2 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

Group feature screening for ultrahigh-dimensional data missing at random

Hanji He1Meini Li2Guangming Deng3,4( )
School of Economics and Finance, South China University of Technology, Guangdong 510006, China
School of Mathematics and Computer Science, Chongqing College of International Business and Economics, Chongqing 401520, China
School of Mathematics and Statistics, Guilin University of Technology, Guangxi 541000, China
Applied Statistics Institute, Guilin University of Technology, Guangxi 541000, China
Show Author Information

Abstract

Statistical inference for missing data is common in data analysis, and there are still widespread cases of missing data in big data. The literature has discussed the practicability of two-stage feature screening with categorical covariates missing at random (IMCSIS). Therefore, we propose group feature screening for ultrahigh-dimensional data with categorical covariates missing at random (GIMCSIS), which can be used to effectively select important features. The proposed method expands the scope of IMCSIS and further improves the performance of classification learning when covariates are missing. Based on the adjusted Pearson chi-square statistics, a two-stage group feature screening method is modeled, and theoretical analysis proves that the proposed method conforms to the sure screening property. In a numerical simulation, GIMCSIS can achieve better finite sample performance under binary and multivariate response variables and multi-classification covariates. The empirical analysis through multiple classification results shows that GIMCSIS is superior to IMCSIS in imbalanced data classification.

CLC number: 62H30, 62R07

References

【1】
【1】
 
 
AIMS Mathematics
Pages 4032-4056

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
He H, Li M, Deng G. Group feature screening for ultrahigh-dimensional data missing at random. AIMS Mathematics, 2024, 9(2): 4032-4056. https://doi.org/10.3934/math.2024197

2

Views

0

Downloads

0

Crossref

0

Web of Science

1

Scopus

Received: 21 November 2023
Revised: 23 December 2023
Accepted: 05 January 2024
Published: 15 February 2024
©2024 the Author(s), licensee AIMS Press.

This is an open access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0)