TY - GEN
T1 - Towards the use of machine learning algorithms to enhance the effectiveness of search strings in secondary studies
AU - Cairo, Leonardo
AU - Monteiro, Miguel P.
AU - Carneiro, Glauco de F.
AU - Brito E Abreu, Fernando
PY - 2019/9/23
Y1 - 2019/9/23
N2 - Devising an appropriate Search String for a secondary study is not a trivial task and identifying suitable keywords has been reported in the literature as a difficulty faced by researchers. A poorly chosen Search String may compromise the quality of the secondary study, by missing relevant studies or leading to overwork in subsequent steps of the secondary study, in case irrelevant studies are selected. In this paper, we propose an approach for the creation and calibration of a Search String. We chose three published systematic literature reviews (SLRs) from Scopus and applied Machine Learning algorithms to create the corresponding Search Strings to be used in the SLRs. Comparison of results obtained with those published in previous SLRs, show an increase of recall of revisions by up to 12%, with no loss of recall. To motivate future studies and replications, the tool implementing the proposed approach is available in a public repository, along with the dataset used in this paper.
AB - Devising an appropriate Search String for a secondary study is not a trivial task and identifying suitable keywords has been reported in the literature as a difficulty faced by researchers. A poorly chosen Search String may compromise the quality of the secondary study, by missing relevant studies or leading to overwork in subsequent steps of the secondary study, in case irrelevant studies are selected. In this paper, we propose an approach for the creation and calibration of a Search String. We chose three published systematic literature reviews (SLRs) from Scopus and applied Machine Learning algorithms to create the corresponding Search Strings to be used in the SLRs. Comparison of results obtained with those published in previous SLRs, show an increase of recall of revisions by up to 12%, with no loss of recall. To motivate future studies and replications, the tool implementing the proposed approach is available in a public repository, along with the dataset used in this paper.
KW - Machine learning
KW - Natural language processing
KW - Secondary studies
UR - https://www.scopus.com/pages/publications/85073189678
U2 - 10.1145/3350768.3350772
DO - 10.1145/3350768.3350772
M3 - Conference contribution
AN - SCOPUS:85073189678
T3 - ACM International Conference Proceeding Series
SP - 22
EP - 26
BT - Proceedings of the 33rd Brazilian Symposium on Software Engineering, SBES 2019
PB - ACM - Association for Computing Machinery
T2 - 33rd Brazilian Symposium on Software Engineering, SBES 2019
Y2 - 23 September 2019 through 27 September 2019
ER -