Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Aligning Large Multimodal Model with Sequential Recommendation via Content-Behavior Guidance

作者:Zihao Wu, Xin Wang, Heng Chang, Hong Chen, Lifeng Sun, Wenwu Zhu · 年份:2025 · DOI:10.1145/3731715.3733273 · 被引用次数:3 · 研究领域:Recommender Systems and Techniques、Topic Modeling、Image Retrieval and Classification Techniques

Large language models (LLMs) have significantly influenced advancements in sequential recommendation. Nevertheless, the integration and alignment of LLMs with sequence recommenders is often underexploited in current research. Existing LLM-based sequential recommenders mostly rely on textual descriptions, neglecting user visual preferences and suffering from LLM hallucination, which can result in suboptimal recommendations. To address these challenges, we propose AlmostRec, a novel framework that incorporates multimodal information, including historical interaction IDs, textual descriptions, and images of items into large foundation models, to facilitate controllable predictions. Instead of employing a textual LLM, AlmostRec utilizes a large multimodal model (LMM) as a backbone, complemented by a content-behavior guidance module to align multimodal information. The framework's ID prediction objective, enhanced via the parameter-efficient LoRA approach, ensures a principal alignment with the sequential recommendation and is not swayed by hallucinations. AlmostRec effectively bridges the gap between large vision-language models with sequential recommenders, offering contextually relevant predictions in multimodal scenarios. Experimental results on real-world datasets demonstrate the superior performance of AlmostRec compared to both traditional and recent LLM-based recommendation approaches.