Abstract
With the rapid adoption of containerized applications, ensuring the security of container images has become a critical concern. Traditional image scanning tools provide a list of vulnerabilities but lack automated risk classification mechanisms that aid in proactive mitigation. This research presents a Machine Learning (ML)-based approach to classify container images into High-Risk and Low-Risk categories using metadata and vulnerability scan results. The dataset was generated by scanning widely used Docker images with Trivy, capturing attributes such as image size, number of packages, file count, executables, and CVE severity levels. Two XGBoost-based classification models were developed. The first model used raw features, achieving an accuracy of 90.91%. Employing the same datasets, the second model achieved 100% accuracy using engineering features, specifically Vuln_Score (Critical + High vulnerabilities) and Pkg_per_MB (package density). The results show that adding domain-specific features improves risk detection accuracy and provides a scalable way to automate security assessments in CI/CD pipelines. This study proposes an effective method for classifying container images and detecting security flaws for different containerized platforms.
Author supplied keywords
Cite
CITATION STYLE
Ugale, S., & Potgantwar, A. (2025). Risk Classification of Docker Container Images Using Machine Learning and Vulnerability Metrics. Engineering, Technology and Applied Science Research, 15(6), 30310–30316. https://doi.org/10.48084/etasr.12627
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.