Abstract
Encrypted traffic has been known to be vulnerable to traffic analysis attacks that exploit the statistical features of encrypted traffic flows, such as packet sizes, timing, and direction, to infer information about the underlying content, which undermines the privacy guarantees of end-to-end encryption. Existing methods such as CNNs lack generalizability, requiring model changes based on the dataset. Furthermore, while state-of-the-art attacks leverage deep learning models to achieve high accuracy, most attacks work under the less realistic closed-set assumption, failing in the open-set setting. Deploying such attacks in practice requires addressing the open-set scenario, which allows the models to filter out target content from other background traffic. The effectiveness of open-set traffic classification largely relies on the model's ability to generalize and accurately extract features from traffic traces, which are essentially sequential data. Concurrently, Large Language Models (LLM) are increasingly becoming popular in modeling sequential data beyond their typical applications in natural language processing. Inspired by this, our work introduces TrafficLLM, a novel traffic analysis attack method that leverages pre-trained LLMs, such as GPT-2 and LLaMA-2-7B, to extract features from network traffic traces with minimal fine-tuning. Using seven existing encrypted traffic datasets, we show that LLMs improve the open-set performance of traffic classification; for instance, our method, TrafficLLM outperforms ET-BERT and CNN-based approaches by 12.7 % and 13.7 % with GPT-2 feature extractor and 17.6 % and 21.5 % with LLaMA-2-7B feature extractor, respectively.
Author supplied keywords
Cite
CITATION STYLE
Ginige, Y., Silva, B., Dahanayaka, T., & Seneviratne, S. (2026). TrafficLLM: LLMs for improved open-set encrypted traffic analysis. Computer Networks, 274. https://doi.org/10.1016/j.comnet.2025.111847
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.