Holistic 3 D Indoor Scene Parsing and Reconstruction from a Single RGB Image

Siyuan Huang,Siyuan Qi,Yixin Zhu,Yinxue Xiao,Yuanlu Xu,Song-Chun Zhu
2018-01-01
Abstract:We propose a computational framework to parse and reconstruct the 3D configuration of an indoor scene from a single RGB image using a stochastic grammar model. We introduce a Holistic Scene Grammar (HSG) to represent the 3D scene structure, which characterizes a joint distribution over the functional and geometric space of indoor scenes. The proposed HSG captures three essential but often latent dimensions of the indoor scenes: i) latent human context, describing the affordance and the functionality of a room arrangement, ii) geometric constraints over the scene configurations, and iii) physical constraints that guarantee physically plausible parsing and reconstruction. We solve this parsing and reconstruction problem in an analysis-bysynthesis fashion, seeking to minimize the differences between the input image and the rendered image generated by our 3D representation, over the space of depth, surface normal, and object segmentation map. The optimal configuration (i.e., parse graph) is inferred using Markov chain Monte Carlo (MCMC), which efficiently traverses through the non-differentiable solution space, jointly optimizing object localization, 3D layout, and hidden human context.
What problem does this paper attempt to address?