MoDiPO: text-to-motion alignment via AI-feedback-driven direct preference optimization
作者:Massimiliano Pappa, Luca Collorone, Giovanni Ficarra, Indro Spinelli, Fabio Galasso · 发表于:Frontiers in Computer Science · 年份:2026 · DOI:10.3389/fcomp.2026.1707808 · 被引用次数:1 · 研究领域:Human Motion and Animation
Diffusion Models have revolutionized the field of human motion generation by offering exceptional generation quality and fine-grained controllability through natural language conditioning. Their inherent stochasticity, that is the ability to generate various outputs from the same input prompt, is key to their success. However, this diversity should not be unrestricted, as it may lead to unlikely generations. Instead, it should be confined within the boundaries of text-aligned and realistic generations. To address this issue, we propose MoDiPO (Motion Diffusion DPO), the first methodology to adapt Diffusion Direct Preference Optimization to align text-to-motion diffusion models. We streamline the laborious and expensive process of gathering human preferences needed in DPO by leveraging AI feedback instead. This enables us to experiment with novel DPO strategies, using both online and offline generated motion-preference pairs. To foster future research we contribute with a motion-preference dataset which we dub Pick-a-Move. We demonstrate, both qualitatively and quantitatively, that our proposed method yields significantly more realistic motions. In particular, MoDiPO achieves statistically significant improvements in Fréchet Inception Distance (FID) of up to 39% on MLD/HumanML3D and consistent gains of 9%–15% across both MLD and MDM on HumanML3D and KIT-ML. Finally, MoDiPO secures a threefold increase in preference from human evaluators compared to the original models' outputs...