Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
The existing cross-modal retrieval algorithms based on metric learning ignore the pose differences and domain differences in cross-modal face retrieval tasks. In addition, these algorithms lack learning of global information in the process of metric learning and construct a large number of redundant triplets. Therefore, a cross-modal common representation generation algorithm based on metric learning was proposed in this paper. The algorithm uses the yaw angle equivariant module compensating for yaw angle differences to obtain the image features with robustness, uses the multi-layer attention mechanism to obtain video features with differentiability, uses global triplets and local triplets to jointly train the cross-modal common representation generation network, so as to improve the consistency and accuracy of metric learning. Then it accelerates the convergence of loss functions through the screening of semi-hard triplets. This study proposed a domain adaption algorithm which combines domain calibration and transfer learning to improve the generalization of common representations. The results of comparative experiments on three face video datasets, namely, PB, YTC, and UMD, demonstrate that the algorithm can improve the accuracy of cross-modal face retrieval, and fine-tuning the cross-modal common representation generation network with few samples can improve the accuracy of cross-modal retrieval using target domain images.
Comments on this article