Abstract
In recent years, the application of the LLM model has played an increasing role in more and more fields, including network security. Some attackers exploit LLM to generate malicious code, craft phishing emails, and analyze software vulnerabilities. It also inspires us to utilize LLM to maintain network security. In previous research on malware detection, feature engineering often relied heavily on expert analysis, making the process both challenging and resource-intensive, especially given the rapid evolution and constant updates of malware. Therefore, we propose a malware detection method for intrinsic semantics. The method first designs an API intrinsic semantic feature encoder, which extracts intrinsic semantic features from API names and Microsoft’s official API definitions based on LLM-based prompt engineering and sentence embedding techniques. Then the API co-occurrence feature encoder is designed, which mines the contextual co-occurrence features of API from API call sequences based on the word2vec. The API semantic features and API co-occurrence features are combined to improve the malware detection performance. Also, it uses TCN- GRU to capture dependencies between API calls. Results on several public datasets show that our method achieves better performance than other methods, and in addition, ablation study results demonstrate the important role of intrinsic semantics in malware detection algorithms.
Author supplied keywords
Cite
CITATION STYLE
Hou, R., Tian, X., & Geng, G. (2025). A malware detection method based on LLM to mine semantics of API. EAI Endorsed Transactions on AI and Robotics, 4. https://doi.org/10.4108/airo.8880
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.