Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Efficient and Global Interaction-Aware Retraining-Free Token Pruning for Vision Transformers

作者:Chenqi Shi, Yi Tian, Hongkun Du, Yuxin Li, Ge Zhang, Guangzhen Yao, Yu Li, Niu Qiang, Zhang Wenxin, Zhenyu Yu, Renda Han · 年份:2026 · DOI:10.1109/icassp55912.2026.11463577 · 研究领域:Advanced Vision and Imaging、Advanced Memory and Neural Computing、Advanced Image and Video Retrieval Techniques

Vision Transformers (ViTs) and their variants have achieved remarkable performance across a wide range of challenging tasks, yet they typically incur high computational cost and long inference latency. Token pruning mitigates this issue by removing redundant tokens, thereby reducing resource consumption and accelerating inference. However, existing methods, especially reinforcement learning (RL)-based approaches, often rely solely on sparse accuracy rewards despite improving mask quality through global interaction modeling. Such a mechanism fails to impose explicit constraints on intermediate feature representations, which can cause structural collapse of the feature manifold after pruning. To address this limitation, we propose EGRT (Efficient and Global Interaction-aware Retraining-Free Token Pruning for ViTs). EGRT consists of two cooperative stages. In the Fisher-Prior State Initialization (FPSI) stage, we construct a high-quality pruning prior from the second-order moment of gradients, substantially reducing the exploration uncertainty of RL in high-dimensional discrete spaces. In the Global Mask Optimization (GMO) stage, we introduce a Manifold-Preserving Reward that explicitly penalizes topological discrepancies of features before and after pruning, compelling the agent to automatically maintain feature consistency during search. Extensive experiments on ViT-L, DeiT-B, DeiT-S, and DeiT-T demonstrate that, without any post-hoc parameter tuning, EGRT achieves a better tr...