Scholay

学术搜索 · AI 审稿 · LaTeX 协作

A 28nm 0.22μJ/Token Memory-Compute-Intensity-Aware CNN-Transformer Accelerator with Hybrid-Attention-Based Layer-Fusion and Cascaded Pruning for Semantic-Segmentation

作者:Pingcheng Dong, Yonghao Tan, Xuejiao Liu, Peng Luo, Yu Liu, Luhong Liang, Yitong Zhou, Di Pang, Manto Yung, Dong Zhang, Xijie Huang, Shih-Yang Liu, Yongkun Wu, Fengshi Tian, Chi-Ying Tsui, Fengbin Tu, Kwang‐Ting Cheng · 年份:2025 · DOI:10.1109/isscc49661.2025.10904499 · 被引用次数:10 · 研究领域:Advanced Memory and Neural Computing、Advanced Neural Network Applications、Ferroelectric and Negative Capacitance Devices

Recently, hybrid models integrating a CNN and a Transformer (ConvFormer), shown in Fig. 23.2.1, have achieved significant advancements in semantic segmentation tasks [1]–[4], which are critical for autonomous driving and embodied intelligence. The CNN enhances the multi-scale feature extraction ability of the Transformer to achieve pixel-level classification, but the large token length (TL) demand of semantic segmentation (> 16K TL) incurs significant computation and memory overheads. Prior NN accelerators [5]–[12] demonstrate that sparse computing and pruning can effectively reduce computation and weight storage, but most of them focus on pure CNN or Transformer models in simpler vision or language-processing tasks (1-4K TL). Moreover, the performance bottlenecks of ConvFormers stem from their memory-intensive Backbone and compute-intensive Segmentation Head (Seg. Head), raising three challenges for hardware acceleration: 1) Conventional sparse attention [5]–[9] fails to buffer the attention feature map (Fmap) on-chip when the TL exceeds 16K, even at 90% sparsity, resulting in massive external memory access (EMA). 2) While Layer-Fusion (LF) [13]–[18] is a common technique to reduce Fmap EMA, it is infeasible to buffer key (K), value (V), and convolution weights on-chip simultaneously. Moreover, different fused attention-convolution layers may cover various vanilla attention (VA) tiles, leading to enormous redundant KV and weight EMA. 3) In the Seg. Head, the Fmap sparsity is...