Publications
Sort:
Issue
Extracting apple planting areas from GF-2 satellite imagery using an improved UNet++
Transactions of the Chinese Society of Agricultural Engineering 2025, 41(23): 125-134
Published: 15 December 2025
Abstract PDF (1.7 MB) Collect
Downloads:1

Accurate extraction of apple planting areas from the high-resolution remote sensing images is often required to optimize the production and industrial layout. This study aims to solve the technical problems with the extraction of apple planting areas from the high-resolution remote sensing images, including spatial fragmentation, spectral confusion, and boundary blurring. To this end, an enhanced MSDAW-UNet++ model was proposed using the UNet++ architecture. A multi-scale dual attention (MSDA) module was incorporated at the key feature fusion nodes, in order to enhance the multi-scale contextual and spatial information of apple planting areas; Meanwhile, a wavelet nested fusion (WNFB) module was embedded at the first feature fusion node of each layer in UNet++. A systematic preprocessing was performed on the GF-2 satellite images during data preparation. A dataset was then constructed to extract the apple planting areas. The MSDA module was then integrated with the multi-scale feature extraction, the multi-head self-attention (MHSA), and positional attention (PSA) mechanisms. Four scales of feature representation were firstly obtained using depthwise separable convolutions. Then, this multi-scale information was input into the MHSA and PSA, respectively. The MHSA mechanism was used to construct the long-range dependencies between apple planting areas in the different regions and combined local and global information by the correlations among input sequence elements. After that, the overall structure of the apple orchard was effectively analyzed after calculation. The conventional feed-forward neural network (FFN) was replaced with an enhanced E-FFN. More efficient feature interaction and multi-scale learning were achieved at the lower computational cost. Furthermore, the local perception of a convolutional neural network (CNN) was integrated with the global modelling strength of transformers, in order to enhance the accuracy and efficiency of apple plantation extraction. The PSA was generated the location-aware attention maps using feature interaction, and then explicitly modeled the geometric constraints among pixels. Continuous energy responses were obtained in the edge region. Spatial continuity was captured to reinforce the correlation between long-distance pixels for the smooth transitions between adjacent pixels. The local features were preserved to prevent the spatial disconnection that caused by environmental complexity. Ultimately, the boundary consistency and regional integrity were improved after semantic segmentation. Finally, the PSA and MHSA mechanisms were combined to produce the output features with the multi-scale contextual and spatial information. The apple planting area was extracted in a complex planting environment. In addition to the MSDA module, the wavelet nested fusion (WNFB) module was specifically designed to combine the wavelet transform convolution (WTConv). The conventional convolution kernels were replaced with the wavelet transform convolution to optimize the semantic segmentation using frequency-space domain synergistic feature extraction. The better performance was obtained to differentiate between spectrally similar features. Experimental test showed that the MSDAW-UNet++ model performed best to extract the apple planting areas, with an F1-score of 96.63% and an IoUz of 90.46%. Compared with the UNet++ benchmark model, the improved model was achieved in the absolute improvements of 3.87 percentage points in the F1-score and 10.07 percentage points in the IoU value, respectively. Compared with classic semantic segmentation models (FCN, UNet and DeepLab v3+), current mainstream remote sensing semantic segmentation models (MCSNet, CMLFormer, and CMTFNet), and UNet derivative models (MAResU-Net, CM-UNet, and UNet3+), the F1-score was improved by 2.35-9.55 percentage points and the IoU by 6.67-19.54 percentage points. Ablation experiments were used to analyze the effectiveness of the multi-scale dual-attention and wavelet nested fusion modules. The MSDA and WNFB modules were effectively extracted the multi-scale contextual and spatial information, as well as frequency-domain features of the apple planting area. A more comprehensive feature expression can be provided for the fine extraction of the apple planting area in a complex planting environment. The findings can offer a valuable reference for the fine extraction from the orchard images using high-resolution remote sensing.

Issue
Semantic segmentation model for agricultural greenhouses in high-resolution remote sensing images using an improved U-Net
Transactions of the Chinese Society of Agricultural Engineering 2025, 41(21): 155-164
Published: 15 November 2025
Abstract PDF (2.1 MB) Collect
Downloads:7

Precision agriculture is often required to accurately and rapidly identify the spatial distribution of the greenhouses in the high-resolution remote sensing imagery. However, there are the complex spectral features, variable spatial patterns, and blurred boundaries in such targets. Particularly, conventional models (such as U-Net and DeepLabv3+) have limited to capture the multi-scale contextual information and spectral confusion between greenhouses and background objects, like roads, buildings, or bare soil. Their segmentation performance has significantly restricted to the dense and heterogeneous agricultural landscapes. In this study, a more accurate and generalizable semantic segmentation was developed to specifically extract the greenhouses under complex environmental conditions. An improved semantic segmentation framework (named MEDNet, Multi-scale Edge-enhanced Dense-skip-connection Network) was also proposed to enhance both feature representation and boundary precision. The modified U-Net architecture was constructed to introduce three components. Firstly, the Dense Cross-layer Skip Connection (DCSC) mechanism was selected to replace the conventional skip connections in U-Net. Multi-level semantic and spatial features were integrated after dense hierarchical fusion. The contextual awareness was improved to reduce the information loss during feature propagation. Secondly, the Edge-Aware Feature Enhancement Module (EAFEM) was implemented to combine the Sobel gradient operators and Swin Transformer attention blocks. The detail edge information was better captured, particularly in cases where the adjacent greenhouses shared the similar textures or overlapping boundaries. Thirdly, the Multi-Scale Spectral Enhancement Module (MS-SEM) was introduced to leverage the strong class separability of near-infrared (NIR) spectral features. Multi-scale dilated convolution was utilized with a channel attention mechanism, in order to highlight the greenhouse-specific spectral responses while suppressing irrelevant background noise. The improved model was trained and then evaluated using two types of high-resolution satellite images: QuickBird and GeoEye-1. Four-band multispectral data was provided with the spatial resolutions of 0.6 and 0.5 m, respectively. There were the preprocessing steps, such as atmospheric correction, image fusion, and manual annotation of greenhouse masks. A balanced dataset was then generated for model training, validation, and testing. A comparison was also conducted with four baseline models—U-Net, DeepLabv3+, MACU-Net, and U-Netformer—on the same datasets. Evaluation metrics included Precision, Recall, Accuracy, F1-score, and Intersection over Union (IoU). Results showed that the MEDNet achieved notably higher performance than before. On QuickBird imagery, the MEDNet was attained an F1-score of 93.11% and an IoU of 84.42%, whereas on GeoEye-1 data, the F1-score and IoU reached 95.59% and 90.67%, respectively. There was the improvements of up to 3.07 percentage points in F1 and 6.52 in IoU over the above baseline models. Ablation experiments were further conducted to isolate the contributions of each architectural module. The DCSC component improved the cross-scale feature, while the EAFEM enhanced the boundary localization to avoid the edge ambiguity. The MS-SEM especially effectively separated the greenhouses from spectrally similar features, such as the urban infrastructure or vegetation. Qualitative evaluations over multiple scene types—including dense greenhouse clusters, sparsely distributed greenhouses, and greenhouses with colorful film covers—demonstrated that the robustness and adaptability of MEDNet were achieved to treat the spatial arrangements and visual conditions. In conclusion, the improved MEDNet model substantially improved the accuracy, robustness, and generalization of the agricultural greenhouse segmentation in the high-resolution remote sensing imagery. Multi-scale spatial features, boundary enhancement, and spectral optimization were integrated for the large-scale and automatic monitoring of facility agriculture. The finding can also provide the valuable technical support to the agricultural land planning, rural revitalization, and food security.

Issue
Named entity recognition in the apple cultivation field based on multi-feature fusion
Transactions of the Chinese Society of Agricultural Engineering 2025, 41(10): 176-185
Published: 30 May 2025
Abstract PDF (948.4 KB) Collect
Downloads:1

Named entity recognition (NER) plays a crucial role in various subsequent tasks, including vertical domain information extraction, knowledge graph construction, intelligent question-answering services, etc. To address the issue of low recognition accuracy in the field of apple cultivation, which was caused by scarce annotated data, single-dimensional character embedding representation, and insufficient ability to mine multi-dimensional features, a Chinese apple cultivation named entity recognition Model (ACNM) based on data augmentation and multi-feature fusion was proposed. Firstly, focusing on the primary production processes in apple cultivation, an Apple Cultivation NER Dataset (ACND) covering 14 entity categories was constructed, and a data augmentation layer was then designed to perform entity-level and sentence-level data enhancement. Secondly, a Multi-feature-layer with a pre-trained model, glyph, radical, and lexicon (MPGRL) was designed to extract and dynamically integrate the multi-dimensional features of the apple cultivation texts, including character, glyph, radical, and lexicon embeddings, and the semantic representation of characters was thus enhanced by incorporating dynamic word representations, visual morphological features of Chinese characters, internal structures of Chinese characters, and lexical boundary information. The BERT with whole word masking (BERT-WWM) pre-trained model was employed to acquire character embeddings, incorporating lexical-level semantic information and mitigating the problem of polysemy. The Vision Transformer model was utilized to obtain glyph embeddings, by modeling and learning the visual features of glyphs. The Mamba model was applied to extract radical embeddings while preserving the radical features that encompass richer semantic information. The SoftLexicon method was adopted to acquire lexical embeddings and enhance lexical boundary information. Thirdly, the receptance weighted key value (RWKV) model framework, which combines the advantages of the Transformer’s parallel training and RNN’s efficient reasoning ability, was adopted as the encoding layer to fully extract the semantic information from the MPGRL, and thus better explore the long-range sequence contextual semantics of the apple cultivation text. Finally, the conditional random field (CRF) was used to learn the constraint relationships between different label sequences, thereby obtaining the optimal label sequence for the apple cultivation NER task. The experimental results showed that the data augmentation technology combined with the multi-feature fusion strategy effectively improved the NER accuracy of the model. The F1 value of the ACNM model on the ACND dataset reached 97.02%, which was 2.93~7.80 percentage points higher than the compared model. This indicated that the ACNM model could efficiently extract the rich semantic information from the MPGRL, ultimately improving the NER accuracy of apple cultivation. The ablation experiment results demonstrated that after removing the data augmentation layer, the MPGRL module, and the RWKV module separately from the ACNM model, the F1 values decreased by 2.11, 3.43, and 0.77 percentage points, respectively, suggesting that each module designed had made a positive contribution to the ACNM model. Compared with the fusion of three types of features adopted in this study, the F1 value decreased by 0.85, 0.33, and 2.02 percentage points respectively after eliminating the glyph feature, radical feature, and vocabulary feature. This indicated that all features had a positive effect on improving the accuracy of entity recognition, the combination of glyph and radical could encode Chinese characters from data of two particle sizes and two modalities, playing a complementary role, the combination of the above three external features could comprehensively model the semantics of apple cultivation text, improving the accuracy of model recognition. Experiments were also conducted on three publicly available datasets, namely CLUENER2020, CCKS2017and Boson, and the F1 values achieved 79.33%, 95.20%, and 83.06%, respectively, which also outperformed other comparative models, indicating that the model had certain generalization ability. This study has practical value for the construction of the apple cultivation knowledge graph and also can provide technical reference for NER research of other crops.

Total 3