Generalizable Egocentric Task Verification via Cross-Modal Hybrid Hypergraph Matching
作者:Xun Jiang, Xing Xu, Zheng Wang, Jingkuan Song, Fumin Shen, Heng Tao Shen · 发表于:IEEE Transactions on Pattern Analysis and Machine Intelligence · 年份:2026 · DOI:10.1109/tpami.2026.3655147 · 被引用次数:2 · 研究领域:Multimodal Machine Learning Applications、Advanced Image and Video Retrieval Techniques、Video Analysis and Summarization
Egocentric Task Verification (ETV) aims to determine if the operation flows of procedural tasks in egocentric videos align with the logic of given rules. Early works adopt the video-based verification paradigm that compares a reference video to the testing video, which limits the flexibility of model deployment. Recent researches incorporate reference textual rules instead of videos, describing the operational logic with natural language, but also raises the challenges of cross-modal heterogeneity and hierarchical misalignment between the two modalities. While previous works mainly address the cross-modal heterogeneity between vision and text modalities, they inevitably suffer from two additional key challenges: (1) Existing methods are mostly developed in synthetic domains, yet have not considered the issues of synthetic-to-realistic generalization challenges in real-world applications. (2) The intricate relations between visual content and textual rule involve multiple matching correlations, indicating high-order matching interactions. To address these issues, we proposed the Generalizable Egocentric Task Verification (GETV), and construct a cross-domain ETV benchmark dataset, EgoCross. It features synthetic-to-real cross-domain evaluation, covering both synthetic datasets for training and realistic datasets for testing, across three different types of tasks. Furthermore, we also propose a novel method for this challenge, termed Cross-modal Hybrid Hypergraph Matching (CHHM)...