Despite their transformative potential, large language models (LLMs) remain underutilized in agriculture due to domain-specific data scarcity and computational constraints. This study presents CAMAGRI-GPT, a parameter-efficient agricultural consultation system that addresses these critical challenges through innovative domain adaptation. A corpus of 2.3 million annotated entries was constructed from raw documents (κ=0.82 agreement, 18 categories) and employed LoRA (r=8) and P-tuning v2 to reduce trainable parameters to 0.2% while maintaining 95.8% performance. The RAG framework with HNSW indexing achieves (87±12) ms retrieval latency, enabling real-time consultation. CAMAGRI-GPT demonstrated over 90.0% accuracy across three representative agricultural tasks (crop management, pest and disease diagnosis, and agricultural Q&A), consistently outperforming GPT-3 and BERT-Agri baselines, p<0.001. Median response latency remained below 2 s across all query categories, meeting field deployment requirements. These results demonstrate that domain-adapted LLMs can effectively deliver expert-level agricultural knowledge to resource-constrained farming communities, offering scalable and sustainable solutions to complement declining traditional extension services.
- Article type
- Year
- Co-author
Open Access
Issue
The fluctuations in the supply, consumption, and prices of agricultural products directly affect market monitoring and early warning systems. With the ongoing transformation of China's agricultural production methods and market system, advancements in data acquisition technologies have led to an explosive growth in agricultural data. However, the complexity of the data, the narrow applicability of existing models, and their limited adaptability still present significant challenges in monitoring and forecasting the interlinked dynamics of multiple agricultural products. The efficient and accurate forecasting of agricultural market trends is critical for timely policy interventions and disaster management, particularly in a country with a rapidly changing agricultural landscape like China. Consequently, there is a pressing need to develop deep learning models that are tailored to the unique characteristics of Chinese agricultural data. These models should enhance the monitoring and early warning capabilities of agricultural markets, thus enabling precise decision-making and effective emergency responses.
An integrated forecasting methodology was proposed based on deep learning techniques, leveraging multi-dimensional agricultural data resources from China. The research introduced several models tailored to different aspects of agricultural market forecasting. For production prediction, a generative adversarial network and residual network collaborative model (GAN-ResNet) was employed. For consumption forecasting, a variational autoencoder and ridge regression (VAE-Ridge) model was used, while price prediction was handled by an Adaptive-Transformer model. A key feature of the study was the adoption of an "offline computing and visualization separation" strategy within the Chinese agricultural monitoring and early warning system (CAMES). This strategy ensures that model training and inference are performed offline, with the results transmitted to the front-end system for visualization using lightweight tools such as ECharts. This approach balances computational complexity with the need for real-time early warnings, allowing for more efficient resource allocation and faster response times. The corn, tomato, and live pig market data used in this study covered production, consumption and price data from 1980 to 2023, providing comprehensive data support for model training.
The deep learning models proposed in this study significantly enhanced the forecasting accuracy for various agricultural products. For instance, the GAN-ResNet model, when used to predict maize yield at the county level, achieved a mean absolute percentage error (MAPE) of 6.58%. The VAE-Ridge model, applied to pig consumption forecasting, achieved a MAPE of 6.28%, while the Adaptive-Transformer model, used for tomato price prediction, results in a MAPE of 2.25%. These results highlighted the effectiveness of deep learning models in handling complex, nonlinear relationships inherent in agricultural data. Additionally, the models demonstrate notable robustness and adaptability when confronted with challenges such as sparse data, seasonal market fluctuations, and heterogeneous data sources. The GAN-ResNet model excels in capturing the nonlinear fluctuations in production data, particularly in response to external factors such as climate conditions. Its capacity to integrate data from diverse sources—including weather data and historical yield data—made it highly effective for production forecasting, especially in regions with varying climatic conditions. The VAE-Ridge model addressed the issue of data sparsity, particularly in the context of consumption data, and provided valuable insights into the underlying relationships between market demand, macroeconomic factors, and seasonal fluctuations. Finally, the Adaptive-Transformer model stand out in price prediction, with its ability to capture both short-term price fluctuations and long-term price trends, even under extreme market conditions.
This study presents a comprehensive deep learning-based forecasting approach for agricultural market monitoring and early warning. The integration of multiple models for production, consumption, and price prediction provides a systematic, effective, and scalable tool for supporting agricultural decision-making. The proposed models demonstrate excellent performance in handling the nonlinearities and seasonal fluctuations characteristic of agricultural markets. Furthermore, the models' ability to process and integrate heterogeneous data sources enhances their predictive power and makes them highly suitable for application in real-world agricultural monitoring systems. Future research will focus on optimizing model parameters, enhancing model adaptability, and expanding the system to incorporate additional agricultural products and more complex market conditions. These improvements will help increase the stability and practical applicability of the system, thus further enhancing its potential for real-time market monitoring and early warning capabilities.
In the context of intensified global climate change and frequent meteorological disasters, exploring the significance of meteorological factors on maize yield and accurately predicting maize yield is crucial for enhancing agricultural production and field management. This paper aims to quantitatively analyze the importance of meteorological factors during various growth stages of maize on yield and to establish a highly accurate and reliable maize meteorological yield stacking ensemble learning estimation model for yield prediction.
Using the HP filter method and moving average method, trend yield models for various counties were determined, and county-level meteorological yields were isolated. Three ensemble learning methods (light gradient boosting machine (LightGBM), Bagging, and Stacking) were employed. By analyzing daily meteorological data and maize yield data over 34 years from 596 county-level administrative regions and meteorological observation stations across 12 provinces in China, three maize meteorological yield prediction models based on different ensemble learning frameworks (LightGBM, Bagging, and Stacking) were established.
The HP filter method as the trend yield model was mainly applicable in the regions of Shaanxi, Henan, Jiangsu, and Anhui. Compared to the HP filter method, more counties were suitable for the moving average method, with most counties having the R2 distribution above 0.8. Based on a 5-year sliding forecast and model accuracy evaluation indicators, the mean absolute percentage error (MAPE) for the three models on maize yield was below 6%. The Stacking model achieved a MAPE of 4.60%, indicating high prediction accuracy and strong generalizability. The results demonstrate that the maize meteorological yield stack-integrated learning prediction model has higher accuracy and stronger robustness. It effectively utilizes the characteristics and advantages of each base learner to improve prediction accuracy, making it the optimal model for predicting maize yield based on meteorological factors. Furthermore, a quantitative analysis of the impact of 27 meteorological factors during the maize growth stages in 12 provinces, using the random forest feature importance score, is of reference value for crop monitoring and field management.
The three ensemble learning methods, especially the stack-integrated learning model (Stacking), can accurately reflect the spatiotemporal distribution changes in maize yield. The stack-integrated learning model for maize yield based on meteorological factors provides a new method for field management and accurate prediction of maize yield.
京公网安备11010802044758号