Abstract
Antimicrobial resistant strains of pathogenic Escherichia coli are a burden on the healthcare system, causing longer hospital stays and increased treatment costs compared to nonresistant strains. With whole genome sequencing almost ubiquitous in the analyses of outbreak and surveillance samples, in silico methods for feature identification can be faster and cheaper than traditional wet-lab methods. In this study, machine learning (ML) classification methods were used to predict antimicrobial resistance (AMR) and identify novel genomic markers of resistance. A total of 4300 E. coli whole genome sequences with laboratory-derived susceptible, intermediate, or resistant (SIR) data for 34 antimicrobials were collected. Three models — gradient boosted decision trees, support vector machines (SVMs), and artificial neural networks (ANNs) —were trained using genome subsequences (k-mers) of length 11 to classify unknown isolates as SIR for each antimicrobial. The models achieved high average accuracies (93.6%, 92.7%, and 92.8%, respectively) for our dataset, outperforming database methods including AM-RFinderPlus (63.9%) and ResFinder (75.7%). Tested on two smaller independent datasets, the models’ average accuracies were 81.6% (XGB), 79.9% (SVM), and 81.2% (ANN), while ResFinder’s average accuracy was 94.7%. An advantage of ML models over database methods is that they can identify novel markers of resistance, which is a key advantage for surveillance and research. As more genomic and AMR data become publicly available, these models are expected to further improve in performance and utility.
Author supplied keywords
Cite
CITATION STYLE
Moat, J., Zovoilis, A., Steinkey, R., Zaheer, R., McAllister, T., & Laing, C. (2025). Machine learning methods to identify markers and predict antimicrobial resistance in Escherichia coli. Canadian Journal of Microbiology, 71, 1–15. https://doi.org/10.1139/cjm-2024-0208
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.