The Dunhuang Chinese manuscripts hold a vital position in the study of Chinese civilization, and character-level annotation is of key significance for document digitization, knowledge mining, and cultural heritage preservation. This paper focuses on the automated recognition and annotation of Dunhuang manuscript images, conducting systematic research in three aspects: dataset construction, model training, and system development. A high-quality annotated dataset covering multiple manuscript volumes-with character-level bounding boxes and label information-was built to serve as a foundational resource for subsequent recognition and analysis tasks. Furthermore, a character-level annotation system integrating image preprocessing, layout analysis, character recognition, and manual proofreading was developed, significantly improving annotation efficiency and accuracy through text-recognition algorithms. The research outcomes have been applied to projects involving the collation and preservation of Dunhuang manuscripts, providing a transferable technical framework and practical experience for the intelligent processing of ancient texts.
Publications
- Article type
- Year
Article type
Year
Open Access
Research Article
Issue
Journal of Beijing University of Chemical Technology (Natural Science Edition) 2025, 52(5): 68-75
Published: 20 September 2025
Downloads:0
Total 1
京公网安备11010802044758号