AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Regular Paper

Exploiting the Community Structure of Fraudulent Keywords for Fraud Detection in Web Search

Network Technology Research Center, Institute of Computing Technology, Chinese Academy of Sciences Beijing 100190, China
University of Chinese Academy of Sciences, Beijing 100049, China
Global Energy Interconnection Research Institute Co., Ltd., Beijing 102209, China
LISTIC Laboratory of Computer Science, Systems, Information and Knowledge Processing, Université Savoie Mont Blanc, Chambéry 73011, France
Computer Network Information Center, Chinese Academy of Sciences, Beijing 100190, China
Show Author Information

Abstract

Internet users heavily rely on web search engines for their intended information. The major revenue of search engines is advertisements (or ads). However, the search advertising suffers from fraud. Fraudsters generate fake traffic which does not reach the intended audience, and increases the cost of the advertisers. Therefore, it is critical to detect fraud in web search. Previous studies solve this problem through fraudster detection (especially bots) by leveraging fraudsters’ unique behaviors. However, they may fail to detect new means of fraud, such as crowdsourcing fraud, since crowd workers behave in part like normal users. To this end, this paper proposes an approach to detecting fraud in web search from the perspective of fraudulent keywords. We begin by using a unique dataset of 150 million web search logs to examine the discriminating features of fraudulent keywords. Specifically, we model the temporal correlation of fraudulent keywords as a graph, which reveals a very well-connected community structure. Next, we design DFW (detection of fraudulent keywords) that mines the temporal correlations between candidate fraudulent keywords and a given list of seeds. In particular, DFW leverages several refinements to filter out non-fraudulent keywords that co-occur with seeds occasionally. The evaluation using the search logs shows that DFW achieves high fraud detection precision (99%) and accuracy (93%). A further analysis reveals several typical temporal evolution patterns of fraudulent keywords and the co-existence of both bots and crowd workers as fraudsters for web search fraud.

Electronic Supplementary Material

Download File(s)
jcst-36-5-1167-Highlights.pdf (111.6 KB)

References

【1】
【1】
 
 
Journal of Computer Science and Technology
Pages 1167-1183

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Yang D-H, Li Z-Y, Wang X-H, et al. Exploiting the Community Structure of Fraudulent Keywords for Fraud Detection in Web Search. Journal of Computer Science and Technology, 2021, 36(5): 1167-1183. https://doi.org/10.1007/s11390-021-0218-2

984

Views

2

Crossref

3

Web of Science

2

Scopus

0

CSCD

Received: 11 December 2019
Accepted: 01 July 2021
Published: 30 September 2021
© Institute of Computing Technology, Chinese Academy of Sciences 2021