Vegetation water use is often required to precisely divide the evapotranspiration (ET) components, and then optimize the carbon-water cycle response to climate change. However, conventional classification cannot fully meet the large-scale production in recent years, due to the complex parameterization or costly isotope techniques. It is difficult to capture the highly nonlinear interactions between environmental drivers and physiological responses. Therefore, this study aims to systematically evaluate the effectiveness and accuracy of the machine learning (ML) models in this partitioning task. The ecological research was also combined to determine the key environmental drivers. A dataset was constructed to fully meet the multiple constraints, such as the total primary productivity (GPP) and conservative surface moisture index (CSWI). Water resource utilization efficiency (WUE) was then predicted using the classic TEA algorithm (using random forest), XGBoost (extreme gradient enhancement), LightGBM (light gradient enhancement machine), and SVM (support vector machine). Then, the transpiration (T) was estimated at each time step. The T outputs of the three carbohydrate-coupled models (CASTANEA, JSBach, and MuSICA) were compared to evaluate the data-driven models. The results show that the machine learning model was effectively estimated T, among which the XGBoost exhibited the better performance over the TEA algorithm among all three carbon-water coupling models. The average reduction in root mean square error (RMSE) reached 18% (P<0.001), and the adjusted coefficient of determination value increased, with most sites higher than 0.85 (P<0.001). The statistical significance test(P<0.001) confirmed that the prediction errors were reduced for the high goodness-of-fit. Furthermore, the CSWI threshold of -0.5mm significantly enhanced the training data and the stability of the model after hyperparameter optimization, especially under the complex, high-latitude, and multi-layered vegetation canopy structures. In contrast, the XGBoost and TEA algorithms were superior to the LightGBM and SVM at most sites, in terms of the T estimation. Especially, LightGBm and SVM were also limited in the complex ecological scenarios, such as the high-dimensional feature spaces, due to the feature expression and the insufficient adaptability of the kernel function. The hierarchical feature interaction with the boost and robust constraints can be expected to effectively analyze the inherent nonlinear dynamics in the carbon-water coupling, indicating the significant advantages of the XGBoost. This finding can provide a powerful verification and technical approach for the dynamic, precise, and mechanical information quantification of the vegetation water consumption. The regional and global climate models can be parameterized to enhance the prediction for the ecosystem responses under future climate scenarios in a sustainable forest.
- Article type
- Year
Net ecosystem exchange (NEE) measurement is very critical to understanding carbon flux in ecosystems. But some gaps are still common in data collection, due to the harsh weather or sensor malfunction. Traditional interpolation can struggle with the long-term missing data, leading to inaccuracies in carbon flux analysis. Long-term missing data can also occur in NEE measurements. In this study, the Adapter-Reverseformer (ARformer) model was proposed to enhance the accuracy of NEE gap-filling, especially for extended periods of data loss. A multi-layer perceptron (MLP) was integrated with the Reverseformer module. The longer data gaps were effectively utilized to leverage both environmental data and the temporal patterns in the NEE. A novel system was established to focus on the non-linear relationships between NEE and environmental factors. The time-dependent trend was characterized by NEE data. The obtained model was tested using the FLUXNET 2015 dataset. Half-hourly carbon flux data was collected from 65 sites across 10 types of land use. Five artificial gap scenarios were generated by randomly removing data for continuous periods of 1, 7, 15, 30, and 90 days. The performance of ARformer was compared with marginal distribution sampling (MDS), random forest (RF), and three advanced deep learning models: DLinear, PatchTST, and iTransformer. The results demonstrated that the ARformer model consistently outperformed the baseline, especially when dealing with long-term missing data. Specifically, the performance of RF decreased significantly, when the missing data spanned 90 days, and MDS failed to reasonably estimate the model. In contrast, the ARformer model maintained a high accuracy, with R2 values ranging from 0.762 to 0.913. The root mean square error (RMSE) ranged between 0.668 and 2.724 μmol/(m2·s), the mean absolute error (MAE) ranged from 0.410 to 1.751 μmol/(m2·s), and bias values remained between −0.024 and 0.067 μmol/(m2·s). The ARformer model demonstrated superior performance across different land-use types, including closed shrublands, deciduous broadleaf forests, evergreen broadleaf forests, evergreen needle-leaf forests, and mixed forests. As such, the ARformer model was used to more effectively capture these vegetation types with the complex relationships between NEE and environmental drivers. Furthermore, it was observed that the time-series deep learning models in general provided the better interpolation for the long-term missing NEE data, with ARformer leading in accuracy. In conclusion, deep learning models, particularly the ARformer model, were highly effective in filling the gaps in NEE data for the various ecosystems. The ARformer model was recommended when the data gaps were extended beyond 30 days. The accuracy of interpolation was also attributed to the temporal dependencies and the relationship between NEE and environmental factors. More reliable NEE data was then obtained to clarify the carbon flux dynamics across different ecosystems. Thus, the ARformer model was represented for the long-term data gaps in NEE measurements.
京公网安备11010802044758号