Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Scene Interpretation by Deep Generative Model Utilizing Information of Backgrounds

作者:Yuya Kobayashi, Masahiro Suzuki, Yutaka Matsuo · 发表于:Transactions of the Japanese Society for Artificial Intelligence · 年份:2023 · DOI:10.1527/tjsai.38-3_e-l35 · 被引用次数:1 · 研究领域:Advanced Image and Video Retrieval Techniques、Robotics and Sensor-Based Localization、Remote Sensing and LiDAR Applications

The ability to understand surrounding environment compositionally by decomposing it into its individual components is important cognitive ability. Human beings decompose arbitral entities into some parts based on its semantics or functionality, and recognize those parts as “object”. Such kind of object recognition ability is fundamental to planning. Recently, researches called “scene interpretation” have been conducted using deep generative models. Those researches build models that are able to recognize environment compositionally. The objective of this paper is to extend scene interpretation methods. Application of existing methods are restricted to simple images, and could not deal with complex images such as real images and heavily textured images. This is because previous works are done in fully-unsupervised manner, and the objective function is just minimizing reconstruction error. Therefore, in this case, models have no clues about objects unlike models leveraging supervised information, or inductive bias. In this research, we propose a method to decompose scenes as intended using minimum auxiliary information to identify objects. We build a model that utilizes background as auxiliary information to separate representation of background and foreground, and then we show our method is able to deal with datasets that are difficult for existing methods.