The existing deep network-based no-reference image quality assessment algorithms for authentic distortions have poor performance in representing the quality of natural scene images, which limited their evaluation accuracy and generalization ability. To solve this problem, this paper proposed a deep neural network based on fuse multi-scale features layer-by-layer (MsFF-Net). Firstly, the pre-trained ResNet-50 was used to extract multi-scale features of the image. Then, a multi-scale features fusion module was proposed, which gradually fused adjacent-scale features layer-by-layer to obtain multi-scale fused features that can accurately represent the image quality. The low-dimensional features were further extracted from the multi-scale fused features to obtain multi-granularity image quality perception features. Finally, regression was performed on the low-dimensional features by using a fully connected network which was adaptively generated by the highest-level features. The simulation results show that MsFF-Net outperforms most of the current methods on authentic distortions databases, and it achieves excellent performances on synthetic distortions databases.
- Article type
- Year
In the prediction-residual reconstruction framework, multi-hypothesis prediction through video temporal correlation is a key step in video compressed sensing reconstruction. Aiming at the problems that the current video compressive sensing multi-hypothesis reconstruction neural network has insufficient prediction accuracy and poor theoretical explanation, this paper proposed a feature-domain multi-hypothesis prediction video compressive sensing reconstruction network (FMH_CVSNet) based on the traditional multi-hypothesis theory. Firstly, a new feature-domain multi-hypothesis prediction module was proposed to enhance the prediction ability of the network by constructing a reasonable motion estimation module and hypothesis weight solving module. Then, a two-stage multi-refe-rence frame motion compensation mode was proposed to adapt the sequence features to construct a better hypothesis set to further improve the prediction accuracy. The simulation results show that FMH_CVSNet achieves better reconstruction performance under all experimental conditions, and the average PSNR is improved by 4.76 and 3.87 dB, respectively, compared with 2 sMHR and VCSNet-2.
Traditional video compression coding methods are widely used. In order to further improve the compression performance, research on deep learning-based video compression coding methods has received increasing attention. Existing deep learning video compression coding methods realize motion compensation based on optical flow, which will produce artifacts during the optical flow alignment process, reducing the accuracy of prediction. This paper proposed a motion estimation idea in the deep feature domain, and designed a corresponding neural network to extract motion information in the deep feature domain. On this basis, it proposed a multi-layer multi-hypothesis prediction motion compensation network. By using the multi-hypothesis prediction module in the deep feature domain, the shallow feature domain and the pixel domain, the accuracy of motion compensation was improved, thereby improving the overall rate-distortion performance. Simulation results show that the inter-frame prediction results of the algorithm in the paper mitigate artifacts and the visual effect is significantly better than optical flow alignment. At the same time, the proposed algorithm achieves better rate-distortion performance compared with traditional H.264 and H.265 methods and single-frame reference methods DVC and DVCpro based on deep learning. Compared with the DCVC method at the forefront of research, the algorithm reduces the coding time by approximately 26.8% while the rate distortion performance is similar. Taking the H.264 encoding result as the benchmark, under the condition of the same bit rate, the decoding quality was improved by 3.73 dB, 4.76 dB and 2.65 dB on HEVC test sequences ClassB, ClassD and ClassE. The simulation experiment results show that, when compressing and coding video sequences, the algorithm proposed in the paper can improve the accuracy of motion compensation prediction frames, reduce the prediction error, shortens the residual signal compression coding code stream and improve the overall rate distortion performance.
Compressed sensing theory can be used to solve the problem of limited computing resources of information source acquisition equipment, but there is uncertainty in the signal reconstruction process. Traditional reconstruction algorithm is difficult to be applied in practice because of its high computational complexity. Recently, the reconstruction algorithm based on deep learning has broken the limitation of traditional algorithms, and has attracted wide attention with its fast reconstruction speed and high quality. Existing deep learning reconstruction algorithms can be divided into two types: “black box” and optimization-based inspired network. Compared with the “black box” network structure, the optimization-inspired deep network is easier to obtain high-precision recovery and more interpretable. However, the existing image compressed sensing reconstruction optimization-inspired networks only learn a single gradient in each optimization phase and has shortcomings such as insufficient use of measured information and difficulty in learning gradients, limiting the improvement of reconstruction performance. In order to make full use of the measurement and reduce the difficulty of gradient learning, the idea of high-dimensional space gradient learning was proposed to achieve more accurate gradient regression. On this basis, this paper proposed Feature-domain Proximal High-Dimensional Gradient Descent (FPHGD) algorithm, and designed a Feature-domain Proximal High-dimensional Gradient Descent Network (FPHGD-Net) to realize this algorithm, so as to obtain high-precision image reconstruction. In addition, three kinds of deep space proximal mapping network structures with different complexity were designed to meet different application. According to the spatial complexity from low to high, the corresponding models are respectively called FPHGD-Net-Tiny, FPHGD-Net and FPHGDNet-Plus. Extensive experiment shows that, compared with OPINE-Net+, the average PSNR of the three proposed models on Set11 increase 1.34, 1.51 and 1.88 dB, and recover richer image details in the reconstruction of visual effects.
The existing video compressive sensing reconstruction network usually uses the optical flow network to achieve pixel domain motion estimation and motion compensation. However, during the reconstruction process, the input of the optical flow network is the estimated frame with poor quality, resulting in inaccurate optical flow. The optical flow-based pixel domain alignment and fusion operation will cause noise accumulation, lead to obvious artificial effects in video reconstruction frames and affect the reconstruction quality. Based on the fact that multi-channel information in the feature space has strong robustness to interference noise, this paper applied the idea of feature space optimization to the design of the video compressive sensing reconstruction neural network, and proposed a feature-space optimization-inspired and flow-guided multi-hypothesis cross-attention network (FOFMCNet). To avoid the image structure destruction caused by the noise in the optical flow when warping the image, the study designed multi-hypothesis motion estimation module guided by optical flow and the motion compensation module based on cross-attention to realize the motion estimation and motion compensation of inter-frame in feature space, so as to make full use of inter-frame correlation to assist non-key frame reconstruction. In order to strengthen the reuse of effective information in the process of feature optimization, improve the learning ability of the network and alleviate the problem of gradient explosion, this paper designed a feature-space optimization-inspired u-shape network (FOUNet) as a sub-network of FOFMCNet. Through the cascade of multiple FOUNets, the FOFMCNet realizes the optimization and reconstruction of non-key frames in the feature space. Experimental results show that the reconstruction results of the proposed algorithm are obviously better than those of the existing video compression sensing algorithms on the classical low-resolution dataset (UCF-101 and QCIF) and new high-resolution dataset (REDS4).
京公网安备11010802044758号