Plant diseases remain a major constraint on crop productivity, requiring timely and accurate diagnostic approaches to secure agricultural yields. While existing automated diagnosis methods primarily rely on image data and achieve notable results, their performance often declines in complex field environments with noise and interference. Multimodal learning provides a promising solution by integrating complementary cues from various data sources. However, the heterogeneity between plant phenotypes and other modalities, such as textual descriptions, poses a significant challenge for effective fusion. To address this issue, we propose PlantIF, a multimodal feature interactive fusion model for plant disease diagnosis based on graph learning. PlantIF comprises three key components: image and text feature extractors, semantic space encoders, and a multimodal feature fusion module. Specifically, we employ pre-trained image and text feature extractors to extract visual and textual features enriched with prior knowledge of plant diseases. Semantic space encoders then map these features into both shared and modality-specific spaces, enabling the capture of cross-modal and unique semantic information. To enhance context understanding, we design a multimodal feature fusion module to process and fuse different modal semantic information, and then extract the spatial dependency between plant phenotype and text semantics through the self-attention graph convolution network. We evaluate PlantIF on a multimodal plant disease dataset with 205,007 images and 410,014 texts, achieving 96.95 % accuracy, 1.49 % higher than existing models. These results demonstrate the potential of multimodal learning in plant disease diagnosis and highlight PlantIF's value in precision agriculture. Codes are available at https://github.com/GZU-SAMLab/PlantIF.
- Article type
- Year
- Co-author
Open Access
Research Article
Issue
Open Access
Research Article
Issue
Segmentation of vegetation remote sensing images can minimize the interference of background, thus achieving efficient monitoring and analysis for vegetation information. The segmentation of vegetation poses a significant challenge due to the inherently complex environmental conditions. Currently, there is a growing trend of using spectral sensing combined with deep learning for field vegetation segmentation to cope with complex environments. However, two major constraints remain: the high cost of equipment required for field spectral data collection; the availability of field datasets is limited and data annotation is time-consuming and labor-intensive. To address these challenges, we propose a weakly supervised approach for field vegetation segmentation by using spectral reconstruction (SR) techniques as the foundation and drawing on the theory of vegetation index (VI). Specifically, to reduce the cost of data acquisition, we propose SRCNet and SRANet based on convolution and attention structure to reconstruct multispectral images of fields, respectively. Then, borrowing from the VI principle, we aggregate the reconstructed data to establish the connection of spectral bands, obtaining more salient vegetation information. Finally, we employ the adaptation strategy to segment the fused feature map using a weakly supervised method, which does not require manual labeling to obtain a field vegetation segmentation result. Our segmentation method can achieve a Mean Intersection over Union (MIoU) of 0.853 on real field datasets, which outperforms the existing methods. In addition, we have open-sourced a dataset of unmanned aerial vehicle (UAV) RGB-multispectral images, comprising 2358 pairs of samples, to improve the richness of remote sensing agricultural data. The code and data are available at https://github.com/GZU-SAMLab/VegSegment_SR, and http://sr-seg.samlab.cn/.
京公网安备11010802044758号