Scholay

学术搜索 · AI 审稿 · LaTeX 协作

BAPTISM: A Robust Framework for Encrypted Malicious Traffic Identification With Low-Quality Training Data

作者:Xiang Luo, Chang Liu, Gang Xiong, Gaopeng Gou, Zhen Li, Junzheng Shi, Li Guo, Binxing Fang · 发表于:IEEE Transactions on Information Forensics and Security · 年份:2026 · DOI:10.1109/TIFS.2025.3648170 · 被引用次数:1 · 研究领域:Computer Science

Machine learning (ML) is highly effective for accurate encrypted malicious traffic identification by using high-quality training data. In fact, obtaining such data is costly and challenging. As a result, many ML-based models are inevitably trained on low-quality data and perform poorly. To enhance performance, some methods utilize various sample selection techniques to choose confident samples for model training. However, they often rely on a single metric for this selection, which restricts their adaptability across diverse datasets and noise conditions. In this paper, we propose a robust framework BAPTISM for identifying encrypted malicious traffic with low-quality training data. Particularly, BAPTISM selects a suitable base model for each task, and trains it with early stopping to generate traffic representation before overfitting occurs. Then, we devise an adaptive metric selection strategy to select confident samples. By employing two metrics (JSD and CSD) to assess the characteristic of traffic representation from distinct perspectives, we find the more proper metric for each class and apply it for confident sample selection. According to the confident samples and selected metric for each class, we develop a label correction tactic which adapts to class nature to improve the quality of training data. Finally, we employ parallel training strategy to train the base model with the corrected data, further mitigating the impact of low-quality data. We conduct experiments acr...