In recent years, large language models (LLMs) and multimodal large models (MLMs) have made significant achievements in natural language processing and multimodal content understanding. However, these general-purpose models have obvious shortcomings when dealing with tasks related to cultural heritage, such as biased understanding of domain-specific terminology, lack of cultural and historical background leading to superficial answers, and knowledge hallucination issues, making it difficult for the results to meet actual needs. In response to these challenges, this paper first proposes a multimodal large model oriented towards the field of cultural heritage: Bogu-Wenjin. This study first designs a semi-automated strategy to construct a large-scale multimodal cultural heritage dataset and forms a multimodal knowledge graph. Using the constructed dataset, the general large model is trained in two stages: image-text alignment and instruction fine-tuning, to adapt to the specific needs of the cultural heritage field. In addition, a knowledge graph is introduced as an auxiliary knowledge base, and the credibility and interpretability of the model in the field of cultural heritage Q & A tasks are effectively improved through graph-text retrieval and relationship retrieval strategies. Experimental results show that Bogu-Wenjin performs excellently in various aspects such as artifact image description, attribute question answering, and relationship question understanding. Compared with general multimodal large models, it significantly improves the ability to understand and answer complex cultural content, with a comprehensive score increase of 21.4%, 53% and 20.6% in artifact image description, artifact attribute questions, and artifact relationship questions respectively over the second-best model.
Publications
- Article type
- Year
Year
Open Access
Issue
Journal of Northwest University (Natural Science Edition) 2025, 55(6): 1267-1284
Published: 25 December 2025
Downloads:24
Total 1
京公网安备11010802044758号