Abstract
With the rapid advancements in display technology, from high-resolution monitors to mobile phones and immersive VR headsets, the digital visual contents including image, video and 3D with high photorealism is becoming an increasingly important medium for people to perceive, interact with, and experience the world around them. The most straightforward approach to obtain the photorealistic contents is by cameras, however, the captured contents (images and videos) will inevitably suffer from degradation or incomplete information, caused by various factors including undesirable content, im-proper preservation environment, spatial-temporal incoherence and artifacts. Besides 2D signals, reconstructing the complicated 3D contents also becomes significantly more challenging when only limited observations are available.In this thesis, we will investigate how to create high-quality and photorealistic visual contents from only imperfect observations. Our approaches will rely on generative models, which could learn the underlying probability distribution that is as close as possible to the natural, clean and high-fidelity images. By harnessing the powerful priors learned by generative models, we have potentials to resolve the challenges across different modalities of visual content. In image domain, the main challenge is learning a conditional generative model requires abundant paired data, which however is very hard to acquire especially for the degraded images like old photos. To avoid the insufficient data issue, generative models like variational autoencoders (VAEs) could effectively narrow down the domain gap between simulated data and the real-world data in latent space, or synthesize novel consistent contents for the missing pixels. When comes to video data, we further consider the complex motion and extra temporal dimension by building a recurrent transformer network based on GAN. The success of 3D reconstruction comes from multi-view supervision. When only sparse views is provided, the generative models will also play a key role to compensate the missing information for high-quality 3D content creation.
The thesis consists of three parts. In the first part of the thesis, we focus on how to generate photorealistic image content based on an input observation with diverse degradations. The first challenge we tackle is old photo restoration. Unlike conventional image restoration tasks that can be solved through supervised learning, the degradation in real photos is complex and the domain gap between synthetic images and real old photos makes a trained network fail to generalize. Therefore, we propose a novel triplet domain translation framework by concurrently leveraging real photos along with massive synthetic image pairs, effectively alleviating the domain gap issues in the compact latent space and enabling the users to recover very high-quality image contents from the old photo collections. Another challenge we solve in the image domain is image completion. We propose ICT, which models the conditional distribution given the masked observation by leveraging generative approaches. ICT significantly improves the generative capability, completion fidelity and diversity by involving bi-directional transformers and the masked language model (MLM) objective, which overcomes the fixed order issue of auto-regressive model.
The human visual system is capable of perceiving not only visual structure and appearance but also motion and dynamics along the temporal dimension. In the second part of the thesis, we aim to generate photorealistic video content from consecutive degraded frames. More specifically, we present a learning-based framework, recurrent transformer network (RTN), to restore heavily degraded old films. Instead of performing frame-wise restoration, our method is based on the hidden knowledge learned from adjacent frames that contain abundant information about the occlusion, which is beneficial to restore challenging artifacts of each frame while ensuring temporal coherency. To better resolve mixed degradation and compensate for the flow estimation error during frame alignment, we propose to leverage more expressive transformer blocks for spatial prediction.
We then consider a much more challenging application, creating immersive and photorealistic 3D content from limited input information. We further propose CAD, a novel learning paradigm for 3D synthesis that utilizes pre-trained diffusion models. Instead of focusing on mode-seeking like classical score distillation, our method directly models the distribution discrepancy between multi-view renderings and diffusion priors in an adversarial manner, which unlocks the generation of high-fidelity and photorealistic 3D content, conditioned on a single image and prompt. Although the volumetric representation in NeRF allows us to reconstruct high-quality 3D contents, its rendering speed is extremely slow, preventing the usage for interactive applications in virtual and augmented reality. We then resolve this problem by proposing the neural duplex radiance field, a novel approach to distill and bake any NeRFs into highly efficient mesh-based neural representations. We conduct extensive experiments to qualitatively and quantitatively evaluate all the improvements we propose. Results indicate that our methods have shown great superiorities over previous methods.
| Date of Award | 26 Jul 2024 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Jing LIAO (Supervisor) |
Cite this
- Standard