Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
The safety hazards associated with hydropower construction are diverse and often occur in complex and variable spatial contexts. Human visual assessments based on experience are prone to cognitive and psychological biases. Existing studies face key limitations, including unimodal models failing to capture cross-modal hazard features, the limited generalization ability of small models, and the poor domain adaptability of general-purpose pre-trained models. To address these challenges, this study establishes the first multimodal image-text dataset for safety hazards in hydropower engineering. By leveraging the Qwen2.5-VL model, we implement efficient domain adaptation through LoRA and instruction tuning. A multimodal large model fine-tuned for the intelligent recognition of safety hazards in hydropower construction was proposed. Comparative experiments across hazard types, modalities, and model architectures reveal that: (1) data imbalance has a limited impact on performance differences across hazard types, (2) textual descriptions generally convey more critical hazard-related information than visual features, and (3) domain-specific fine-tuning and effective multimodal fusion are identified as key factors in enhancing hazard recognition performance.
The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/), which permits use, distribution and reproduction in any medium, provided the original work is properly cited.
Comments on this article