Multi-Objective Recommendation via Multivariate Policy Learning
作者:Olivier Jeunen, Jatin Mandav, Ivan Potapov, N. Agarwal, Sourabh Vaid, Wenzhe Shi, Aleksei Ustimenko · 年份:2024 · DOI:10.1145/3640457.3688132 · 被引用次数:8 · 研究领域:Recommender Systems and Techniques、Topic Modeling、Advanced Bandit Algorithms Research
Real-world recommender systems often need to balance multiple objectives when deciding which recommendations to present to users. These include behavioural signals (e.g. clicks, shares, dwell time), as well as broader objectives (e.g. diversity, fairness). Scalarisation methods are commonly used to handle this balancing task, where a weighted average of per-objective reward signals determines the final score used for ranking. Naturally, how these weights are computed exactly, is key to success for any online platform.