Unsupervised software repositories mining and its application to code search. (21st November 2019)
- Record Type:
- Journal Article
- Title:
- Unsupervised software repositories mining and its application to code search. (21st November 2019)
- Main Title:
- Unsupervised software repositories mining and its application to code search
- Authors:
- Hu, Gang
Peng, Min
Zhang, Yihan
Xie, Qianqian
Gao, Wang
Yuan, Mengting - Other Names:
- Bishop Judith guestEditor.
Cooper Kendra M.L. guestEditor.
Sharp Helen guestEditor.
Whalen Michael guestEditor. - Abstract:
- Summary: Software repositories are crucial resources for many software tasks, including code retrieval and annotation. Programming forums provide questions and answers (Q&A) from software developers, containing abundant code‐description posts for exchanging knowledge about programming issues. However, most posts provide personal opinions of users that are often not adequately confirmed or outdated. Mining software repositories in such open and unrestricted forums is challenging. Since the posts can be arbitrary and noisy, it is difficult to get unified labels for supervised noise elimination. Different from existing mining approaches, this paper proposes Code‐Description Mining Framework (CodeMF), an unsupervised framework to eliminate noisy posts and extract high quality software repositories from programming forums. CodeMF treats all social features of the posts as discrete‐time signals for kernel principal component analysis and further performs wavelet transform feature fusion to find the delicate changes (noises in temporal signals). We conduct comprehensive experiments on StackOverflow. Experimental results demonstrate that CodeMF can effectively reduce running time and improve precision via mining high‐quality software repositories for various programming languages, especially for the large‐scale codebases. To further illustrate the effect of CodeMF applied in software tasks, we introduce it to improve the performance of query‐expansion code search. Meanwhile, for SQLSummary: Software repositories are crucial resources for many software tasks, including code retrieval and annotation. Programming forums provide questions and answers (Q&A) from software developers, containing abundant code‐description posts for exchanging knowledge about programming issues. However, most posts provide personal opinions of users that are often not adequately confirmed or outdated. Mining software repositories in such open and unrestricted forums is challenging. Since the posts can be arbitrary and noisy, it is difficult to get unified labels for supervised noise elimination. Different from existing mining approaches, this paper proposes Code‐Description Mining Framework (CodeMF), an unsupervised framework to eliminate noisy posts and extract high quality software repositories from programming forums. CodeMF treats all social features of the posts as discrete‐time signals for kernel principal component analysis and further performs wavelet transform feature fusion to find the delicate changes (noises in temporal signals). We conduct comprehensive experiments on StackOverflow. Experimental results demonstrate that CodeMF can effectively reduce running time and improve precision via mining high‐quality software repositories for various programming languages, especially for the large‐scale codebases. To further illustrate the effect of CodeMF applied in software tasks, we introduce it to improve the performance of query‐expansion code search. Meanwhile, for SQL and C# programs, compared to the state‐of‐the‐art query‐expansion method QECK, the improvement of QECKCodeMF is 2% and 6% on Recall@10, and 4% and 14% on mean reciprocal rank, respectively. … (more)
- Is Part Of:
- Software, practice & experience. Volume 50:Number 3(2020)
- Journal:
- Software, practice & experience
- Issue:
- Volume 50:Number 3(2020)
- Issue Display:
- Volume 50, Issue 3 (2020)
- Year:
- 2020
- Volume:
- 50
- Issue:
- 3
- Issue Sort Value:
- 2020-0050-0003-0000
- Page Start:
- 299
- Page End:
- 322
- Publication Date:
- 2019-11-21
- Subjects:
- code search -- feature fusion -- software repositories -- wavelet transformation
Computer software -- Periodicals
Computer programming -- Periodicals
Computer programs -- Periodicals
005.3 - Journal URLs:
- http://onlinelibrary.wiley.com/ ↗
- DOI:
- 10.1002/spe.2760 ↗
- Languages:
- English
- ISSNs:
- 0038-0644
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 8321.453000
British Library DSC - BLDSS-3PM
British Library STI - ELD Digital store - Ingest File:
- 12658.xml