Skip to main navigation Skip to search Skip to main content

Physically Grounded Novel View Synthesis: From Geometric Consistency to Generative Priors

Student thesis: Doctoral Thesis

Abstract

Novel View Synthesis (NVS), which aims to generate photorealistic images of a scene from arbitrary viewpoints, serves as a cornerstone for emerging technologies such as virtual and augmented reality, autonomous navigation, digital twins, and immersive content creation. Recent advances in neural rendering and implicit 3D representations have significantly expanded the capabilities of NVS, enabling the synthesis of highly detailed and complex visual scenes that surpass the limits of traditional geometry-based methods. Despite these advances, existing approaches still struggle to jointly achieve accurate 3D geometry reconstruction, robust dynamic modeling, and strong generalization across diverse real-world scenarios—three factors crucial for practical deployment. To address these challenges, this thesis presents a progressive exploration of NVS, moving from static scene reconstruction to dynamic scene modeling, and ultimately toward a unified and generalizable framework capable of handling both static and dynamic 3D scenes.

We begin with the foundational setting of view synthesis from sparse multi-view inputs in static scenes, which represents the base case for real-world 3D reconstruction. In Chapter 3, we introduce PCVS, a deep learning–based framework that learns a locally unified 3D point-cloud representation from sparse inputs. The model constructs partial point clouds via depth-guided projections and adaptively fuses them within local neighborhoods to form a coherent global structure. A geometry-guided image restoration module further refines the rendered outputs by recovering high-frequency textures and filling occluded regions, leading to photorealistic and geometrically consistent results. Experiments on three benchmark datasets demonstrate that our method substantially improves synthesis quality while preserving fine visual details compared with state-of-the-art view synthesis approaches.

Since monocular video capture is the most common form of real-world dynamic data, the next step focuses on dynamic view synthesis—specifically, generating novel views from arbitrary viewpoints using a monocular video of a dynamic scene captured by a moving camera. The main challenge lies in accurately modeling dynamic objects using only a limited number of 2D frames, each associated with different timestamps and viewpoints. Existing methods often rely on precomputed optical flow or depth maps from off-the-shelf models, which introduce errors and ambiguity when lifting 2D information into 3D. To address this, Chapter 4 presents DecouplingNeRF, a dynamic scene modeling approach that disentangles object motion and camera motion within a Neural Radiance Field (NeRF) formulation. The motion of dynamic objects is separated into object and camera components, each regularized by our proposed unsupervised surface-consistency and patch-based multi-view constraints. The former enforces temporal consistency of 3D geometric surfaces, while the latter ensures appearance coherence across viewpoints. This fine-grained motion formulation reduces the network’s learning burden, enabling it to produce novel views with higher visual quality and more accurate scene flow and depth estimates than existing methods that rely on explicit supervision.

While the first two approaches rely on explicit 3D scene reconstruction, such geometry-based methods inherently struggle to infer unseen regions from the available views and often fail to generalize without strong assumptions, such as dense input coverage or a specific camera trajectory. To overcome these limitations, Chapter 5 introduces NVS-Solver, a training-free paradigm for both static and dynamic NVS that leverages large pretrained video diffusion models. The method adaptively modulates the diffusion sampling process using information from given views, enabling the generation of high-fidelity results from single or multiple images of static scenes as well as monocular videos of dynamic scenes. Building on our theoretical formulation of the diffusion process, we iteratively modulate the score function with scene priors derived from warped input views to guide the diffusion trajectory. Moreover, by analyzing the estimation error, we adaptively adjust the modulation according to view poses and diffusion steps. Extensive evaluations on both static and dynamic benchmarks demonstrate that our method outperforms state-of-the-art alternatives both quantitatively and qualitatively.

Overall, this thesis establishes a coherent research trajectory that bridges geometry-based learning, motion disentanglement, and generative modeling for novel view synthesis. By progressively advancing from static to dynamic scene understanding and ultimately unifying both within a single framework, it pushes NVS toward more flexible, data-efficient, and physically grounded neural rendering—laying a foundation for future developments in immersive visual computing.
Date of Award5 Mar 2026
Original languageEnglish
Awarding Institution
  • City University of Hong Kong
SupervisorJunhui HOU (Supervisor)

Cite this

'