As a key individual feature in forensic investigation and biometric recognition, footprint images are highly susceptible to diverse environmental factors during acquisition, often accompanied by complex noise and image quality degradation. To address composite noise commonly present in footprint images, this paper proposes an enhanced dual-branch cyclic denoising network for high-fidelity image restoration and texture structure reconstruction. The overall network comprises two generators and two discriminators, with the generator comprising two synergistically optimized branches: a denoising mapping branch and a color correction branch. The denoising mapping branch incorporates an Enhanced Multi-Scale Structure Block (EMSB) to strengthen structural modeling and texture recovery capabilities. By integrating multi-scale convolutions, depthwise separable convolutions, and multi-attention mechanisms, this branch effectively enhances feature representation in texture-sensitive regions. Simultaneously, the color correction branch employs an adaptive Color Consistency Module (CCM), which extracts color features via multi-scale residual convolutions and performs channel-wise normalization and residual fusion in the RGB space to suppress color deviation in the generated images. Furthermore, a multi-level structural perception loss function is designed, combining pixel-level accuracy with structural similarity to guide the network in recovering details while improving overall perceptual quality. Experimental evaluations conducted on the self-built footprint dataset, FSD-Real, demonstrate that the proposed method achieves a Peak Signal-to-Noise Ratio (PSNR) of 30.3 dB and a Structural Similarity Index (SSIM) of 0.926, significantly outperforming existing mainstream methods. Moreover, the method exhibits superior denoising performance and detail preservation in terms of subjective visual quality, validating its application potential in real-world footprint image processing tasks.
- Article type
- Year
- Co-author
Aiming at the limited feature extraction capability and inefficient quantization constraint mechanism of existing hashing methods, a deep multi-scale attention hashing network was proposed for large-scale image retrieval. The whole network was composed of a main branch and an object branch. In the main branch, two modules of multi-scale attention localization and saliency region extraction were added to effectively localize and extract saliency regions of images, and the results were fed into the object branch to learn more detailed features. Subsequently, the multi-granularity features learned by two branches were fused to perform binary hash coding. In addition, a triplet quantization constraint was introduced to reduce quantization error while maintaining the similarity relationship between sample pairs. In order to verify the effectiveness of the proposed method, extensive experiments were carried out on two benchmark datasets. Experimental results show that the proposed method outperforms most existing hashing retrieval methods.
Human body mesh reconstruction (HMR) has wide applications in human-computer interaction, virtual/augmented reality, and other fields. In order to further improve the accuracy of human body pose and shape estimation in image-based human body mesh reconstruction, this study proposed a parametric human body mesh reconstruction network based on hybrid inverse kinematics and global consistency deep convolutional neural network, called GloCoNet. To enhance the network’s global consistency and long-range dependencies, a Global Consistency Booster (GCB) module was designed on top of the feature extraction network. It can enhance the model’s perception and expression capabilities of global information, and allow the model to adaptively adjust the feature map weights of different channels and spatial positions. Furthermore, a multi-head attention mechanism was introduced to capture the model’s long-range dependencies globally, helping the model better capture key relationships and patterns when dealing with long-term dependencies, and modeling global contextual information to enrich the diversity of feature subspaces. Meanwhile, the network adopts a hybrid inverse kinematics approach to bridge the gap between human body mesh estimation and 3D human joint estimation, ultimately improving the accuracy of human 3D pose and shape estimation. Experimental results show that the GloCoNet model significantly outperforms previous mainstream methods with an average per joint position error of 51.3 mm on the publicly available Human3.6M dataset.
京公网安备11010802044758号