Sort:
Research Article Issue
Impacts of Preprocessing Methods and Machine Learning Algorithms on Glass Material Property Prediction: Comparative Study and Optimization Strategies
Journal of the Chinese Ceramic Society 2025, 53(10): 2870-2881
Published: 12 September 2025
Abstract PDF (10.1 MB) Collect
Downloads:0
Introduction

The existing research on modeling glass properties has some limitations in critical modeling stages, thus hindering predictive accuracy and interpretability. Primarily, systematic data pre-processing is often neglected, i.e., most literature relies solely on basic data cleaning (e.g., removing outliers and duplicates), lacking thorough investigation into optimal missing value imputation methods and the crucial sequence of imputation versus feature standardization. Secondly, hyperparameter optimization strategies remain simplistic. As some advanced techniques like random search or Bayesian optimization are introduced, studies generally overlook a critical adaptation of tuning methods to specific algorithm characteristics and a lack comparative analysis of hyperparameter optimization approaches across different models. This potentially restricts the performance of hyperparameter-sensitive models like Support Vector Machines (SVM). To address these challenges, this study was to propose a comprehensive, optimized framework for glass property prediction.

Methods

This research established a fully optimized framework for glass property prediction. The methodology comprised four integrated stages, i.e., 1) Systematic Data Preprocessing:For dataset A, the framework rigorously evaluated the combinatorial effects of three missing value imputation strategies (i.e., Median Imputation, Polynomial Imputation, Random Forest Imputation) and two standardization sequences (i.e., Imputation followed by Standardization; Standardization followed by Imputation); 2) Predictive Model Construction and Evaluation:Six distinct machine learning algorithms, which were chosen for their varied inductive biases (i.e., Random Forest-RF, XGBoost-XGB, Classification and Regression Trees-CART, K-Nearest Neighbors-KNN, Support Vector Machines-SVM, Artificial Neural Networks-ANN), were employed. The model predictive efficacy was rigorously quantified using a triple evaluation metric (i.e., Mean Squared Error–MSE, Mean Absolute Error–MAE, Coefficient of Determination-R2). The model generalization was validated by a novel dataset (A); 3) Interpretable Analysis:The SHAP (SHapley Additive exPlanations) framework was applied to elucidate the contribution paths of key glass components towards the predicted properties; and 4) Hyperparameter Optimization Strategy Analysis: For representative models SVM and RF, a comparative experiment was conducted by three hyperparameter tuning methods, i.e., Random Search, Grid Search, and Bayesian Optimization. This stage was to establish adaptable criteria matching hyperparameter tuning strategies to model architectures.

Results and Discussion

The results show that the sequence of standardization and imputation significantly affects model performance metrics, with a substantially larger impact observed on MSE and MAE (up to 49.64% variation), compared to R2. Conversely, the choice of imputation method exerts a greater influence on R2 (up to 63.68%) under conditions of high missing data rates (i.e., >10%). The optimal preprocessing is standardization-first followed by Random Forest Imputation. The SHAP analysis reveals that preprocessing sequence does not alter the fundamental trends in feature contributions, but significantly changes the numerical distribution morphology of SHAP values, thus introducing systematic bias. Imputation method, particularly at high missing rates (i.e., >40%), directly affects the ranking of feature importance derived from SHAP.

The comparative performance of the six algorithms, evaluated across metrics, follows an order of RF ≈ XGB > KNN ≈ ANN > SVM≥ CART. The RF and XGB demonstrate a superior performance, leveraging their ensemble learning advantages. The predictive accuracy of model (as measured by MSE, MAE, and R2) is robust to varying missing data rates. However, the interpretability of model predictions via SHAP becomes less reliable at high missing rates.

The SHAP analysis indicates distinct governing components for different properties, sometimes diverging from conventional understanding. While B2O3 dominates the Glass Transition Temperature (Tg), Mg content is primarily governed by SiO2, and followed by B2O3. Na2O is a dominant contributor to both Density and Refractive Index. The primary contributors to the thermal expansion coefficient (TEC) are temperature-range dependent. Na2O followed by K2O dominates at 20–300 ℃, whereas SiO2 followed by Na2O dominates at 20–400 ℃.

The use of hyperparameter tuning method proves critical, especially for hyperparameter-sensitive models. The Bayesian optimization can obtain the most substantial improvements for the lower-accuracy SVM model, achieving a maximum R2 improvement of 76.56%. In contrast, its impact on the high-performing ensemble algorithm RF is marginal, indicating that alternative, potentially more explorative methods like Genetic Algorithms (GA) can be more suitable for such models. This highlights a necessity for model-specific hyperparameter optimization strategies.

Conclusions

This study proposed and validated a comprehensive optimization framework for glass property prediction, addressing key limitations in data preprocessing and hyperparameter tuning prevalent in the existing literature. The systematic evaluation demonstrated that preprocessing choices, particularly the sequence of operations and imputation method under high missing data scenarios, had profound and quantifiable impacts on both predictive accuracy (MSE, MAE, R2) and the reliability of interpretability frameworks like SHAP. Among the evaluated algorithms, RF and XGB consistently outperformed others. Critically, the SHAP analysis provided insights into composition–property relationships, thus revealing temperature–dependent dominance for TEC and challenging assumptions regarding Mg's primary drivers. Finally, the hyperparameter optimization experiments established that tuning strategy efficacy could be highly model-dependent. The Bayesian optimization offered significant gains for sensitive models like SVM but was less effective for robust ensembles like RF, necessitating tailored approaches. This framework could provide a validated, end-to-end methodology for enhancing the accuracy, reliability, and interpretability of predictive models in glass science and potentially other material domains.

Research Article Issue
Homogenizing Mechanism of Silica Glass via Data Mining of Raman Spectra
Journal of the Chinese Ceramic Society 2025, 53(10): 2830-2840
Published: 12 September 2025
Abstract PDF (1.4 MB) Collect
Downloads:1
Introduction

Quartz glass is a type of inorganic amorphous material prepared by ultra-clean high-temperature melting process with high-purity natural quartz or synthetic silicon compounds as raw materials, having an outstanding physical and chemical property. This material possesses superior optical performance, extreme thermal stability, ultra-low thermal conductivity, and superior dielectric properties. In the development of optoelectronic technology industry in China, quartz glass with its unique synergy of physical and chemical parameters becomes a strategic basic material in the field of optoelectronic devices. The feature size of devices is in a hundred-micron scale based on the iterative evolution of microelectromechanical systems and wafer-level packaging technologies, further demanding a deeper analysis of the surface structure of quartz glass.

Methods

The glass was firstly prepared in a deposition furnace, and then heated in a high-temperature homogenization furnace at 1800 ℃ for 2 h. Afterwards, the glass was cooled to room temperature and discharged. The glass before and after homogenization was cut and polished to obtain glass samples with the sizes of 200 mm×200 mm×5 mm. Starting from the first point position in the upper left corner, some points were taken at intervals of 20 mm in all the directions (i.e., up, down, left, and right) for the analysis of Raman spectroscopy. In addition, the PCA algorithm was utilized for data dimensionality reduction and visualization, and the results of the changes in peak center and half-peak width were also analyzed.

Results and discussion

Before homogenization, the standard deviations of LO_FWHM and D2_FWHM are 15.43 cm-1 and 11.73 cm-1 respectively, which are relatively larger than those of other indicators. The fluctuation range of the two sets of data is also larger, with the mean values of 114.13 cm-1 and 47.30 cm-1, respectively. The standard deviation of D1_center is only 0.18 cm-1, indicating that the data for this indicator are concentrated, with small differences between each data point. The difference between the maximum and minimum values is also similar. For instance, the maximum value of LO_FWHM is 178.79 cm-1 and the minimum value is 100.28 cm-1; the maximum value of D2_FWHM is 115.58 cm-1 and the minimum value is 42.77 cm-1. This further indicates that the dispersion before homogenization of the two sets of data is large.

After homogenization, the standard deviation of D2_center is 0.39 cm-1, which is decreased by 79%, compared to that before homogenization. The standard deviation of D2_FWHM is 0.82 cm-1, which is decreased by 93%, compared to that before homogenization. The standard deviation of LO_FWHM is 8.03 cm-1, which is decreased by 48%, compared to that before homogenization. The data fluctuations between different positions significantly reduce, and the standard deviation statistics for each indicator also show a certain degree of reduction. Among them, the influence of D2 peak is dominant, and the influence of D1 peak is subordinate. The D1 peak and D2 peak are respectively related to the tetra-siloxane ring and tri-siloxane ring in the structure. This indicates that on the same sample surface, the homogenization process has a dominant impact on the tri-siloxane ring in the silicon-oxygen structure. The tri-siloxane ring is expected in quartz glass because their Si—O—Si angle is less than the lowest energy (most likely) angle in the glass. The estimation of the energy required to generate such rings is more likely to reach the distribution of the D2 tri-siloxane ring. Also, the silicon—oxygen bond angles of all rings tend to be averaged, and the overall is more normalized.

Conclusions

The Raman spectroscopy could precisely characterize the microscopic structural differences of quartz glass before and after homogenization via detecting the vibration modes of atoms in the silicon-oxygen network (i.e., Si—O—Si bond angles and ring structure defects). For the analysis of structural uniformity, some parameters such as peak center dispersion and half-width were used to quantify the structural uniformity of the glass surface and bulk phase, and to identify high-strain areas (i.e., the regions with concentrated trimer ring defects). In the study of defect mechanisms, the D1 peak (tetramer ring) and D2 peak (trimer ring) defect characteristic peaks were utilized to reveal the migration and recombination laws of defects during homogenization, and to clarify the influence of thermal stress release on structural relaxation. Finally, combined with the data mining technology of the PCA algorithm, a correlation between the Raman parameters and homogenization processes was established, thus providing a theoretical support for precisely regulating the microscopic structure of quartz glass.

Total 2