Abstract
Software reuse has proven to be an effective strategy for developers to significantly increase software quality, reduce costs and increase the effectiveness of software development. Research in software reuse typically addresses two main hurdles: reduce the time and effort required to identify reusable candidates, and avoid selecting low-quality software components that may lead to higher cost of development (i.e., solving bugs, errors, refactoring). Inherently, human judgment falls short in the aspect of reliability and effectiveness. Hence this paper investigates the applicability of Machine Learning (ML) algorithms in assessing software reuse. We collected more than 32k open-source projects and employed GitHub fork as the ground truth to its reuse. We developed ML classification pipelines based on both internal and external software metrics to perform software reuse prediction. Our best-performing ML classification model achieved an accuracy of 86%, outperforming existing research in prediction performance and data coverage. Subsequently, we leverage our results by identifying key software characteristics that make software highly reusable. Our results show that size-related metrics (i.e., number of setters, methods, attributes) are the most impactful in contributing to the reuse of the software.
Author supplied keywords
Cite
CITATION STYLE
Yeow, M. Y. H., Chong, C. Y., & Lim, M. K. (2022). On the application of machine learning models to assess and predict software reusability. In MaLTeSQuE 2022 - Proceedings of the 6th International Workshop on Machine Learning Techniques for Software Quality Evaluation, co-located with ESEC/FSE 2022 (pp. 17–22). Association for Computing Machinery, Inc. https://doi.org/10.1145/3549034.3561177
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.