Abstract
When looking at an image, instead of a group of pixels and a 2D plane, we humans can perceive more of the semantic meaning and 3D understanding of the scene behind this 2D image. This is achieved by joint reasoning of different representations from the image. Similarly, the computer should also go beyond plain 2D pixels to reveal the 3D scene underlying. In this thesis, we investigate several approaches to leverage synergy representations extracted from the image for scene understanding, in a joint reasoning manner.To well model the stereo vision mechanism of human perception, we first study the 3D information extraction from stereo images. As captured images in the real world can inevitably contain noises, we propose to jointly tackle the problems of 3D prediction and noise removal. Specifically, the depth information and clean images are recovered simultaneously. The joint addressing approach is shown to be more effective than each individual method and the sequential pipeline. The proposed approach is then applied to applications such as novel view synthesis and digital refocusing to show its effectiveness.
Beyond the stereo vision, humans can still well-perceive the 3D scene from a single image. Hence we further investigate the scene understanding from a single image. As high-quality clean image is a fundamental assumption for most computer vision tasks including monocular image processing and understanding, we first propose to deal with a more general image restoration problem beyond noise removal. We show that different types of corruptions can be eliminated at the same time in a single model, by a data-driven approach that learns a common latent space sharing among different corrupted images. We also demonstrate the transferability of the proposed model to deal with other restoration problems like super-resolution, rain removal, depth enhancement, etc.
Finally, we study the problem of 3D scene understanding from monocular image. In addition to the pixel intensity, 2D image is also embedded with higher level semantic meanings. We find that when inferring 3D information from single image, semantic meanings of the objects in the scene have auxiliary effect. On the contrary, the 3D geometry of the scene also plays a synergy role while predicting the semantic labels. As a result, we propose to jointly address the tasks of depth estimation and semantic labeling of the scene, with mutually improved performance. Therefore, given a single image, our model is able to glean 3D information and semantic meanings simultaneously. Above all, our work demonstrates that when understanding the 3D scene from 2D images, either stereo or monocular, by jointly reasoning synergy representations, computer can be endowed with the ability to perceive tremendous amount of information underlying. We hope that this thesis could inspire others to further explore the synergy representations behind the 2D image and better leverage them towards understanding the 3D scene.
| Date of Award | 6 Sept 2018 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Rynson W H LAU (Supervisor) & Qingxiong YANG (Supervisor) |
Cite this
- Standard