AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (2.2 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese | Open Access

Bogu-Wenjin: A cultural heritage domain multimodal large model enhanced by knowledge graphs

Wanqing ZHAO1,2( )Chaoyang XU1Zhiwei XIE1Shaobo ZHANG1,2Xiaodan ZHANG1,2Jinye PENG1,2
School of Electronic Information 〔School of Artifical Intelligence〕, Northwest University, Xi'an 710127, China
Shaanxi Key Laboratory of Higher Education Institution of Generative ArtificialIntelligence and Mixed Reality, Xi'an 710127, China
Show Author Information

Abstract

In recent years, large language models (LLMs) and multimodal large models (MLMs) have made significant achievements in natural language processing and multimodal content understanding. However, these general-purpose models have obvious shortcomings when dealing with tasks related to cultural heritage, such as biased understanding of domain-specific terminology, lack of cultural and historical background leading to superficial answers, and knowledge hallucination issues, making it difficult for the results to meet actual needs. In response to these challenges, this paper first proposes a multimodal large model oriented towards the field of cultural heritage: Bogu-Wenjin. This study first designs a semi-automated strategy to construct a large-scale multimodal cultural heritage dataset and forms a multimodal knowledge graph. Using the constructed dataset, the general large model is trained in two stages: image-text alignment and instruction fine-tuning, to adapt to the specific needs of the cultural heritage field. In addition, a knowledge graph is introduced as an auxiliary knowledge base, and the credibility and interpretability of the model in the field of cultural heritage Q & A tasks are effectively improved through graph-text retrieval and relationship retrieval strategies. Experimental results show that Bogu-Wenjin performs excellently in various aspects such as artifact image description, attribute question answering, and relationship question understanding. Compared with general multimodal large models, it significantly improves the ability to understand and answer complex cultural content, with a comprehensive score increase of 21.4%, 53% and 20.6% in artifact image description, artifact attribute questions, and artifact relationship questions respectively over the second-best model.

CLC number: TP391.4

References

【1】
【1】
 
 
Journal of Northwest University (Natural Science Edition)
Pages 1267-1284

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
ZHAO W, XU C, XIE Z, et al. Bogu-Wenjin: A cultural heritage domain multimodal large model enhanced by knowledge graphs. Journal of Northwest University (Natural Science Edition), 2025, 55(6): 1267-1284. https://doi.org/10.16152/j.cnki.xdxbzr.2025-06-006

1080

Views

22

Downloads

0

Crossref

0

CSCD

Received: 07 July 2025
Published: 25 December 2025
© The Editorial Department of Journal of Northwest University (Natural Science Edition)2025.

This is an open access article under the CC BY-NC-ND 4.0 license (https://creativecommons.org/licenses/by-nc-nd/4.0/).