Skip to main navigation Skip to search Skip to main content

Latent 3D Scene Graph with Aligned Visual-Language Perception for Object-Goal Navigation

  • Jianwei Zhang
  • , Kang Zhou
  • , Junjie Yang
  • , Xiaojian Li*
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Object-goal navigation requires agents to accurately locate and navigate to specified target objects in complex indoor environments. Existing methods primarily rely on single visual observations or simple semantic matching, making it difficult to effectively understand 3D scene relationships between target objects, which leads to inefficient path planning during navigation. Furthermore, the feature discrepancy between visual and language modalities poses challenges for accurate target identification. To address these issues, we propose Latent 3D Scene Graph with Aligned Visual-Language Perception (LSGAP), a novel approach that explicitly models 3D scene relationships between objects through depth perception and latent 3D scene graph construction. Meanwhile, we leverage multimodal pretraining strategies to achieve effective visual-language feature alignment, enhancing the model's target recognition capabilities. This end-to-end architectural design improves spatial cognition ability and enables more efficient navigation decision-making. We conduct extensive experimental evaluations in the AI2THOR interactive environment to validate the effectiveness of our navigator. © 2025 IEEE.
Original languageEnglish
Title of host publication2025 3rd International Conference on Mechatronics, Control and Robotics (ICMCR 2025)
PublisherIEEE
Pages63-68
ISBN (Electronic)979-8-3315-4454-6, 979-8-3315-4453-9
ISBN (Print)979-8-3315-4455-3
DOIs
Publication statusPublished - 2025
Event2025 3rd International Conference on Mechatronics, Control and Robotics (ICMCR 2025)
- , Singapore
Duration: 14 Feb 202516 Feb 2025

Publication series

NameInternational Conference on Mechatronics, Control and Robotics, ICMCR

Conference

Conference2025 3rd International Conference on Mechatronics, Control and Robotics (ICMCR 2025)
PlaceSingapore
Period14/02/2516/02/25

Research Keywords

  • Scene graph
  • Visual navigation
  • VisualLanguage alignment

Fingerprint

Dive into the research topics of 'Latent 3D Scene Graph with Aligned Visual-Language Perception for Object-Goal Navigation'. Together they form a unique fingerprint.

Cite this