Multimodal Deep Learning for Stage Classification of Head and Neck Cancer Using Masked Autoencoders and Vision Transformers with Attention-Based Fusion
作者:A. Turki, Ossama S. Alshabrawy, W. Woo · 发表于:Cancers · 年份:2025 · DOI:10.3390/cancers17132115 · 研究领域:Medicine
Simple Summary Head and neck squamous cell carcinoma (HNSCC) is a common and deadly form of cancer. Doctors use a system known as AJCC staging to determine how advanced the cancer is, which helps guide treatment. However, current staging methods mostly rely on simple anatomical observations and limited clinical information, which may not fully capture the complexity of the disease. This study introduces a new computer-based approach that combines detailed medical images (CT scans) and patient clinical data using advanced artificial intelligence techniques. The method uses a special kind of deep learning called Vision Transformers, combined with masked autoencoders, to learn important features from the images without the need for many labelled examples. It also applies attention mechanisms to focus on the most relevant parts of the images and the clinical data. The study tested this approach on two public datasets and showed that it could accurately predict cancer stages better than existing methods. Importantly, different attention models worked better depending on the amount and balance of available data. This work demonstrates the potential of combining imaging and clinical information through AI to improve cancer staging, which could help doctors make better decisions and ultimately benefit patient care.