Abstract
When conducting a systematic review, screening the vast body of literature to identify the small set of relevant studies is a labour-intensive and error-prone process. Although there is an increasing number of fully automated tools for screening, their performance is suboptimal and varies substantially across review topic areas. Many of these tools are only trained on small datasets, and most are not tested on a wide range of review topic areas. This study presents two systematic review datasets compiled from more than 8600 systematic reviews and more than 540000 abstracts covering 51 research topic areas in health and medical research. These datasets are the largest of their kinds to date. We demonstrate their utility in training and evaluating language models for title and abstract screening. Our dataset includes detailed metadata of each review, including title, background, objectives and selection criteria. We demonstrated that a small language model trained on this dataset with additional metadata has excellent performance with an average recall above 95% and specificity over 70% across a wide range of review topic areas. Future research can build on our dataset to further improve the performance of fully automated tools for systematic review title and abstract screening.
Cite
CITATION STYLE
Chan, G. C. K., He, E., Leung, J., & Verspoor, K. (2025). A comprehensive systematic review dataset is a rich resource for training and evaluation of AI systems for title and abstract screening. Research Synthesis Methods, 16(2), 308–322. https://doi.org/10.1017/rsm.2025.1
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.