Abstract
Contextual information has been shown to be effective
in helping solve various image understanding tasks. Previous works have focused on the extraction of contextual
information from an image and use it to infer the properties
of some object(s) in the image. In this paper, we consider
an inverse problem of how to hallucinate missing contextual information from the properties of a few standalone
objects. We refer to it as scene context prediction. This
problem is difficult as it requires an extensive knowledge
of complex and diverse relationships among different objects in natural scenes. We propose a convolutional neural
network, which takes as input the properties (i.e., category,
shape, and position) of a few standalone objects to predict
an object-level scene layout that compactly encodes the semantics and structure of the scene context where the given
objects are. Our quantitative experiments and user studies show that our model can generate more plausible scene
context than the baseline approach. We demonstrate that
our model allows for the synthesis of realistic scene images
from just partial scene layouts and internally learns useful
features for scene recognition.
| Original language | English |
|---|---|
| Title of host publication | Proceedings - 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019) |
| Publisher | IEEE |
| ISBN (Electronic) | 978-1-7281-3293-8 |
| DOIs | |
| Publication status | Published - Jun 2019 |
| Event | The 32nd meeting of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019) - Long Beach, CA, California, United States Duration: 16 Jun 2019 → 20 Jun 2019 http://cvpr2019.thecvf.com/program/main_conference |
Publication series
| Name | Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition |
|---|---|
| ISSN (Print) | 1063-6919 |
Conference
| Conference | The 32nd meeting of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019) |
|---|---|
| Place | United States |
| City | California |
| Period | 16/06/19 → 20/06/19 |
| Internet address |
Bibliographical note
Research Unit(s) information for this publication is provided by the author(s) concerned.Research Keywords
- Image and Video Synthesis
- Scene Analysis and Understanding
Fingerprint
Dive into the research topics of 'Tell Me Where I Am: Object-level Scene Context Prediction'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver