AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (12.2 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

ViT-DualAtt: An efficient pornographic image classification method based on Vision Transformer with dual attention

Zengyu Cai1Liusen Xu2Jianwei Zhang2,3( )Yuan Feng4Liang Zhu1Fangmei Liu1
School of Computer Science and Technology, Zhengzhou University of Light Industry, Zhengzhou 450003, China
School of Software Engineering, Zhengzhou University of Light Industry, Zhengzhou 450003, China
Research Institute of Industrial Technology, Zhengzhou University of Light Industry, Zhengzhou 450003, China
School of Elechonic Information, Zhengzhou University of Light Industry, Zhengzhou 450003, China
Show Author Information

Abstract

Pornographic images not only pollute the internet environment, but also potentially harm societal values and the mental health of young people. Therefore, accurately classifying and filtering pornographic images is crucial to maintaining the safety of the online community. In this paper, we propose a novel pornographic image classification model named ViT-DualAtt. The model adopts a CNN-Transformer hierarchical structure, combining the strengths of Convolutional Neural Networks (CNNs) and Transformers to effectively capture and integrate both local and global features, thereby enhancing feature representation accuracy and diversity. Moreover, the model integrates multi-head attention and convolutional block attention mechanisms to further improve classification accuracy. Experiments were conducted using the nsfw_data_scrapper dataset publicly available on GitHub by data scientist Alexander Kim. Our results demonstrated that ViT-DualAtt achieved a classification accuracy of 97.2% ± 0.1% in pornographic image classification tasks, outperforming the current state-of-the-art model (RepVGG-SimAM) by 2.7%. Furthermore, the model achieves a pornographic image miss rate of only 1.6%, significantly reducing the risk of pornographic image dissemination on internet platforms.

References

【1】
【1】
 
 
Electronic Research Archive
Pages 6698-6716

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Cai Z, Xu L, Zhang J, et al. ViT-DualAtt: An efficient pornographic image classification method based on Vision Transformer with dual attention. Electronic Research Archive, 2024, 32(12): 6698-6716. https://doi.org/10.3934/era.2024313

98

Views

1

Downloads

3

Crossref

2

Web of Science

2

Scopus

Received: 30 July 2024
Revised: 28 November 2024
Accepted: 05 December 2024
Published: 15 December 2024
©2024 the Author(s), licensee AIMS Press.

This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0)