Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Foundation Models for Multimodal MRI Synthesis with Language Guidance

作者:Mahmut Yurt, Xiaozhi Cao, Zihan Zhou, Kawin Setsompop, Shreyas Vasanawala, John Pauly · 年份:2025 · DOI:10.1109/iccvw69036.2025.00704 · 被引用次数:1 · 研究领域:Speech and dialogue systems、Speech Recognition and Synthesis、Face recognition and analysis

Generalizing across diverse imaging tasks and modalities with minimal supervision remains a core challenge in$3 D$medical imaging, especially in MRI, underscoring the need for robust foundation models. We present FoundSyn, a language-guided foundation model designed to address this challenge by enabling controllable and flexible synthesis for multimodal MRI. FoundSyn leverages a single-step latent diffusion model, conditioned jointly on image and text embeddings, to translate between source and target modalities under language guidance. To support fast adaptation, we incorporate low-rank adaptation, allowing efficient fine-tuning with limited data. Pretrained on the large-scale IXI dataset, FoundSyn demonstrates high-fidelity synthesis and strong generalization across both dataset and modality shifts, adapting effectively with only a few examples. These results position FoundSyn as a versatile framework, demonstrating the potential of language-guided models for flexible and generalizable medical image synthesis.