Scholay

学术搜索 · AI 审稿 · LaTeX 协作

SMGNFORMER: Fusion Mamba‐graph transformer network for human pose estimation

作者:Yi Li, Zan Wang, Weiran Niu · 发表于:IET Computer Vision · 年份:2024 · DOI:10.1049/cvi2.12339 · 被引用次数:5 · 研究领域:Human Pose and Action Recognition、Hand Gesture Recognition Systems、Video Surveillance and Tracking Methods

Abstract In the field of 3D human pose estimation (HPE), many deep learning algorithms overlook the topological relationships between 2D keypoints, resulting in imprecise regression of 3D coordinates and a notable decline in estimation performance. To address this limitation, this paper proposes a novel approach to 3D HPE, termed the Spatial Mamba Graph Convolutional Neural Network (GCN) Former (SMGNFormer). The proposed method utilises the Mamba architecture to extract spatial information from 2D keypoints and integrates GCNs with multi‐head attention mechanisms to build a relational graph of 2D keypoints across a global receptive field. The outputs are subsequently processed by a Time‐Frequency Feature Fusion Transformer to estimate 3D human poses. SMGNFormer demonstrates superior estimation performance on the Human3.6M dataset and real‐world video data compared to most Transformer‐based algorithms. Moreover, the proposed method achieves a training speed comparable to PoseFormerv2, providing a clear advantage over other methods in its category.