Current picture inpainting techniques use auxiliary structural information prediction to fill realistic patches, however erroneous priors can result in unrealistic structures and blurry textures. Meanwhile, existing methods only focus on the relationship between the original image and the inpainted image, and do not fully utilize the information of the damaged image. To address the above problems, an end-to-end transformer face image inpainting network is proposed, which utilizes semantic segmentation and edge texture information to guide the inpainting process. The main inpainting network includes one RGB inpainting branch and two auxiliary branches for semantic segmentation and edge texture. A set of large kernel convolutional context bottleneck (LKCCB) modules is designed in the encoder to increase the effective receptive field and better contextual reasoning. In order to capture distant contextual information, a nested dynamic auxiliary normalization multi-head attention (NDAN-MHA) module is proposed, which contains a dynamic auxiliary normalization (DAN) module that can dynamically integrate the structural features of the three branches to enrich semantic consistency. Furthermore, a contrastive regularization (CR) network is proposed to stabilize and improve the training of the network to generate more realistic inpainted images. The CelebA-HQ and FFHQ datasets were used for both qualitative and quantitative trials. The findings demonstrate that the suggested method performs better than the comparative methods in both subjective and objective measures and that it can reasonably restore huge, irregularly occluded face photos.
Publications
- Article type
- Year
Year
Issue
Journal of Beijing University of Aeronautics and Astronautics 2026, 52(6): 2194-2207
Published: 11 December 2024
Downloads:0
Total 1
京公网安备11010802044758号