Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
Feature selection (FS) in huge data sets is a critical aspect of machine learning that involves choosing the most relevant features. It plays a significant role in improving the model's performance, reducing overfitting, and enhancing the interpretability. In this paper, we construct an automatic modification of the basic FS techniques such as Lasso, Deep Neural Network (DNN), Random Forest (RF) and a Principal Component Analysis (PCA), based on the K-means clustering method and the Silhouette score method, instead of visualization or threshold based methods based on background knowledge. Additionally, the construction of two hybrid methods is proposed, the purpose of which is to exploit the advantages offered by a number of feature seledtion methods: the first is the score method to leverage multiple types of methods; and the second is the refinement method to enhance the outcomes of one method by adapting them to another method. Moreover, to evaluate the efficiency of the FS method, a linear regression and a DNN nonlinear regression are employed to minimize the dependency of the choice of the regression. Through numerical tests, we show that the automatic modification of the conventional methods can generate a convenient way to set the criterion. Additionally, based on the results that are derived from both the linear regression and the DNN regression, the hybrid FS techniques can more accurately perform in both the linear and nonlinear regressions without any dependency on the data.
This is an open access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0)
Comments on this article