AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Regular Paper

Detecting Duplicate Contributions in Pull-Based Model Combining Textual and Change Similarities

Key Laboratory of Parallel and Distributed Computing, College of Computer, National University of Defense Technology Changsha 410073, China
Laboratory of Software Engineering for Complex Systems, College of Computer, National University of Defense Technology, Changsha 410073, China
Show Author Information

Abstract

Communication and coordination between open source software (OSS) developers who do not work physically in the same location have always been the challenging issues. The pull-based development model, as the state-of-the-art collaborative development mechanism, provides high openness and transparency to improve the visibility of contributors’ work. However, duplicate contributions may still be submitted by more than one contributor to solve the same problem due to the parallel and uncoordinated nature of this model. If not detected in time, duplicate pull-requests can cause contributors and reviewers to waste time and energy on redundant work. In this paper, we propose an approach combining textual and change similarities to automatically detect duplicate contributions in the pull-based model at submission time. For a new-arriving contribution, we first compute textual similarity and change similarity between it and other existing contributions. And then our method returns a list of candidate duplicate contributions that are most similar to the new contribution in terms of the combined textual and change similarity. The evaluation shows that 83.4% of the duplicates can be found in average when we use the combined textual and change similarity compared with 54.8% using only textual similarity and 78.2% using only change similarity.

Electronic Supplementary Material

Download File(s)
jcst-36-1-191-Highlights.pdf (688.1 KB)

References

【1】
【1】
 
 
Journal of Computer Science and Technology
Pages 191-206

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Li Z-X, Yu Y, Wang T, et al. Detecting Duplicate Contributions in Pull-Based Model Combining Textual and Change Similarities. Journal of Computer Science and Technology, 2021, 36(1): 191-206. https://doi.org/10.1007/s11390-020-9935-1

1015

Views

13

Crossref

13

Web of Science

14

Scopus

0

CSCD

Received: 15 August 2019
Accepted: 03 January 2021
Published: 05 January 2021
© Institute of Computing Technology, Chinese Academy of Sciences 2021