Spatial Induction of 2D Generative Priors
posted on 4 August, 2026


Abstract: Pre-trained 2D generative models, trained on large-scale image data, implicitly encode rich knowledge of the 3D world in their weights, including object geometry, depth distribution, occlusion relationships, and material regularities. However, there remain gaps between these spatial priors and 3D vision tasks at three levels: representation, objective, and process, making direct transfer difficult. This talk focuses on how to systematically induce spatial intelligence from 2D generative priors. It proposes decomposing the induction mechanism into three independent dimensions: at the representation level, designing bridge representations compatible with 2D priors to enable cross-domain transfer; at the objective level, converting geometric correctness into differentiable optimization signals; and at the process level, replacing stochastic sampling with deterministic pathways to better match the requirements of perception tasks. The talk will analyze representative works in 3D object generation, camera-controllable video generation, depth and normal estimation, and related tasks, demonstrating the effectiveness of these three types of induction mechanisms. It will also discuss the boundaries and frontier directions of transferring 2D generative priors toward spatial intelligence