Abstract
This study presents a multi-modal machine learning framework for automated CVSS severity prediction, addressing scalability and consistency challenges in vulnerability assessment. We integrate semantic code representations from GraphCodeBERT with TF-IDF features, processed through class-weighted gradient boosting, ordinal regression (Ordinal Logistic Regression and CORAL), and cost-sensitive models to respect severity ordering and minimize false negatives for critical threats. Evaluating 36 model configurations, our optimal GraphCodeBERT with Code_TFIDF_Fusion and LightGBM achieves 66.6% accuracy, 78.0% F1-score for CRITICAL vulnerabilities, and 76.0% recall, reducing manual reviews by over three-quarters. Cost-sensitive models further enhance safety by prioritizing CRITICAL vulnerability detection. A Mean Absolute Error of 0.647 ensures errors are mostly between adjacent severity levels, with 45% overestimating severity for conservative safety. Graph structural features and multi-modal fusion improve performance, establishing a robust foundation for efficient, security-oriented vulnerability management.
Author supplied keywords
Cite
CITATION STYLE
Lim, S. C., & Kim, J. C. (2025). Multi-Modal Machine Learning for Automated CVSS Severity Prediction for Conservative Vulnerability Assessment. IEEE Access, 13, 189344–189357. https://doi.org/10.1109/ACCESS.2025.3627919
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.