Access to safe drinking water is a fundamental determinant of global health. The presence of contaminated water affects the citizens’ health. Per- and polyfluoroalkyl substances (PFAS) are often referred to as forever chemicals. They pose a persistent and growing threat to drinking water. In the literature, machine learning methods are used to identify the forever chemicals in water. However, traditional methods are not efficient and scalable. Thus, to solve this issue. This study develops a large-scale machine-learning framework for PFAS risk screening in US public water systems. The proposed framework incorporates data ingestion, preprocessing, and feature engineering. We have used SMOTE for correcting imbalanced data. We performed experimentation and also evaluated our ensemble-based framework integrating Gradient boosting, bagging, and meta-learning strategies. The proposed framework achieves a maximum ROC-AUC of 0.9574, with the best-performing stacking ensemble achieving a precision of 0.75, a recall of 0.68, and an F1-score of 0.71. The simulation results show that the proposed ensemble learning framework is useful for screening and identifying water systems.
- Article type
- Year
- Co-author
Open Access
Article
Issue
Open Access
Article
Issue
The emergence of digital networks and the wide adoption of information on internet platforms have given rise to threats against users’ private information. Many intruders actively seek such private data either for sale or other inappropriate purposes. Similarly, national and international organizations have country-level and company-level private information that could be accessed by different network attacks. Therefore, the need for a Network Intruder Detection System (NIDS) becomes essential for protecting these networks and organizations. In the evolution of NIDS, Artificial Intelligence (AI) assisted tools and methods have been widely adopted to provide effective solutions. However, the development of NIDS still faces challenges at the dataset and machine learning levels, such as large deviations in numeric features, the presence of numerous irrelevant categorical features resulting in reduced cardinality, and class imbalance in multiclass-level data. To address these challenges and offer a unified solution to NIDS development, this study proposes a novel framework that preprocesses datasets and applies a box-cox transformation to linearly transform the numeric features and bring them into closer alignment. Cardinality reduction was applied to categorical features through the binning method. Subsequently, the class imbalance dataset was addressed using the adaptive synthetic sampling data generation method. Finally, the preprocessed, refined, and oversampled feature set was divided into training and test sets with an 80–20 ratio, and two experiments were conducted. In Experiment 1, the binary classification was executed using four machine learning classifiers, with the extra trees classifier achieving the highest accuracy of 97.23% and an AUC of 0.9961. In Experiment 2, multiclass classification was performed, and the extra trees classifier emerged as the most effective, achieving an accuracy of 81.27% and an AUC of 0.97. The results were evaluated based on training, testing, and total time, and a comparative analysis with state-of-the-art studies proved the robustness and significance of the applied methods in developing a timely and precision-efficient solution to NIDS.
京公网安备11010802044758号