AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Regular Paper

NR-CLIP: CLIP-Guided Multimodal News Recommendation via Multi-View Learning

School of Artificial Intelligence and Computer Science, North China University of Technology, Beijing 100144, China
Department of Computer Science and Technology, Tsinghua University, Beijing 100084, China
College of Intelligence and Computing, Tianjin University, Tianjin 300350, China
School of Computer Science, Shanghai Jiao Tong University, Shanghai 200240, China
College of Computer Science and Technology, Guizhou University, Guiyang 550025, China
Show Author Information

Abstract

In the era of social media, the evolution of news has diversified its format, incorporating texts, images, and videos. However, the majority of news recommendation methods focus solely on text data, overlooking the substantial role of news images. This paper introduces a news recommendation method based on CLIP (Contrastive Language-Image Pretraining), NR-CLIP, which is a CLIP-guided algorithm with enhanced multimodal news recommendation via multi-view learning. Specifically, our method employs a CLIP encoder to embed textual and visual information into the same feature space of neural networks, where a unified news textual representation is learned by treating titles, categories, subcategories, and bodies as different views of news. In addition, feature enhancements are applied to fully fuse textual and visual features. Finally, click history and user representations are embedded to predict the click probability of candidate news. Extensive comparisons and in-depth analyses with state-of-the-art news recommendation methods have been presented on V-MIND (Visual Microsoft News Dataset), which provides the visual information on the basis of the classic MIND (Microsoft News Dataset), demonstrating that our proposed method effectively improves the performance of news recommendation.

Electronic Supplementary Material

Download File(s)
JCST-2503-15356Highlights.pdf (193.7 KB)

References

【1】
【1】
 
 
Journal of Computer Science and Technology
Pages 896-909

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Guo Y, Hu S-T, Yu M-J, et al. NR-CLIP: CLIP-Guided Multimodal News Recommendation via Multi-View Learning. Journal of Computer Science and Technology, 2026, 41(3): 896-909. https://doi.org/10.1007/s11390-025-5356-5

5

Views

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 12 March 2025
Accepted: 13 January 2026
Published: 01 May 2026
© Institute of Computing Technology, Chinese Academy of Sciences 2026