Despite their transformative potential, large language models (LLMs) remain underutilized in agriculture due to domain-specific data scarcity and computational constraints. This study presents CAMAGRI-GPT, a parameter-efficient agricultural consultation system that addresses these critical challenges through innovative domain adaptation. A corpus of 2.3 million annotated entries was constructed from raw documents (κ=0.82 agreement, 18 categories) and employed LoRA (r=8) and P-tuning v2 to reduce trainable parameters to 0.2% while maintaining 95.8% performance. The RAG framework with HNSW indexing achieves (87±12) ms retrieval latency, enabling real-time consultation. CAMAGRI-GPT demonstrated over 90.0% accuracy across three representative agricultural tasks (crop management, pest and disease diagnosis, and agricultural Q&A), consistently outperforming GPT-3 and BERT-Agri baselines, p<0.001. Median response latency remained below 2 s across all query categories, meeting field deployment requirements. These results demonstrate that domain-adapted LLMs can effectively deliver expert-level agricultural knowledge to resource-constrained farming communities, offering scalable and sustainable solutions to complement declining traditional extension services.
- Article type
- Year
- Co-author
Open Access
Issue
Accurate prediction of soybean demand is of profound practical significance for safeguarding national food security, optimizing industrial decision-making, and responding to fluctuations in international trade. Traditional soybean demand forecasting methods are plagued by inadequacies such as limited capacity to excavate data dimensionality and multivariate interactive features, insufficient ability to capture nonlinear relationships under the coupling of multi-dimensional dynamic factors, and challenges in model interpretability and domain adaptability. These limitations render them incapable of effectively supporting accurate prediction and interpretable analysis of China's soybean demand. When the temporal fusion transformers (TFT) model is applied to forecast China's soybean demand, it exhibits certain constraints in aspects like feature interaction layers and attention weight allocation. Consequently, there is an urgent need to explore a forecasting method based on the improved TFT model to enhance the accuracy and interpretability of soybean demand prediction.
Drawing on relevant studies, this research applied the deep learning-based TFT model to China's soybean demand forecasting and proposed the MA-TFT (improved TFT model based on MDFI and AAWO) model, which was enhanced through multi-layer dynamic feature interaction (MDFI) and adaptive attention weight optimization (AAWO). Firstly, a dataset for analyzing China's soybean demand, covering eight dimensions: consumption, production, trade, inventory, market, economy, policy, and international factors, was collated. This dataset, encompassing 4652 relevant indicators spanning from 1980 to 2024, was subjected to data cleaning, transformation, augmentation, and feature engineering. The training, validation, and test sets for the model were constructed using the rolling window method. Secondly, based on the architecture of the TFT model for China's soybean demand forecasting, a multi-layer dynamic feature interaction module and an adaptive attention weight optimization strategy were designed. Additionally, the model's loss function, training strategy, and Bayesian hyperparameter tuning method were formulated, and the model performance evaluation metrics were determined. Subsequently, experiments were designed to compare the prediction performance of the MA-TFT model with that of the autoregressive integrated moving average model (ARIMA), the long short-term memory (LSTM) model, and the original TFT model. Ablation experiments on the MDFI and AAWO modules were conducted separately. The SHapley Additive exPlanations (SHAP) tool was employed for interpretability analysis to identify key feature variables influencing China's soybean demand and their interaction relationships. Error analysis was performed between the predicted and actual values of China's historical soybean demand, and a comparative analysis of the predicted soybean demand in China from 2025 to 2034 was carried out.
The mean squared error (MSE) and mean absolute percentage error (MAPE) of the MA-TFT model were 0.036 and 5.89%, respectively, with a coefficient of determination R2 of 0.91, all of which outperformed those of the comparative models, namely ARIMA (1, 1, 1), LSTM, and TFT. Compared with the benchmark TFT model, the root mean square error (RMSE) and MAPE of the MA-TFT model decreased cumulatively by 21.84% and 3.44%, respectively. These results indicated that the MA-TFT model, as an improved version of TFT, could capture complex relationships between features and enhance prediction performance and accuracy. Interpretability analysis using the SHAP tool revealed that the MA-TFT model exhibited high stability in explaining key feature variables affecting China's soybean demand. It was projected that China's soybean demand would reach 117.99 million tons, 110.33 million tons, and 113.78 million tons in 2025, 2030, and 2034, respectively.
The MA-TFT model, developed by improving the TFT model, provides an innovative solution to address the practical issues of insufficient accuracy and poor interpretability in existing soybean demand forecasting methods. It also offers valuable references for method optimization and application in time series forecasting of other bulk agricultural products.
京公网安备11010802044758号