Abstract
Object-goal navigation requires agents to accurately locate and navigate to specified target objects in complex indoor environments. Existing methods primarily rely on single visual observations or simple semantic matching, making it difficult to effectively understand 3D scene relationships between target objects, which leads to inefficient path planning during navigation. Furthermore, the feature discrepancy between visual and language modalities poses challenges for accurate target identification. To address these issues, we propose Latent 3D Scene Graph with Aligned Visual-Language Perception (LSGAP), a novel approach that explicitly models 3D scene relationships between objects through depth perception and latent 3D scene graph construction. Meanwhile, we leverage multimodal pretraining strategies to achieve effective visual-language feature alignment, enhancing the model's target recognition capabilities. This end-to-end architectural design improves spatial cognition ability and enables more efficient navigation decision-making. We conduct extensive experimental evaluations in the AI2THOR interactive environment to validate the effectiveness of our navigator. © 2025 IEEE.
| Original language | English |
|---|---|
| Title of host publication | 2025 3rd International Conference on Mechatronics, Control and Robotics (ICMCR 2025) |
| Publisher | IEEE |
| Pages | 63-68 |
| ISBN (Electronic) | 979-8-3315-4454-6, 979-8-3315-4453-9 |
| ISBN (Print) | 979-8-3315-4455-3 |
| DOIs | |
| Publication status | Published - 2025 |
| Event | 2025 3rd International Conference on Mechatronics, Control and Robotics (ICMCR 2025) - , Singapore Duration: 14 Feb 2025 → 16 Feb 2025 |
Publication series
| Name | International Conference on Mechatronics, Control and Robotics, ICMCR |
|---|
Conference
| Conference | 2025 3rd International Conference on Mechatronics, Control and Robotics (ICMCR 2025) |
|---|---|
| Place | Singapore |
| Period | 14/02/25 → 16/02/25 |
Research Keywords
- Scene graph
- Visual navigation
- VisualLanguage alignment
Fingerprint
Dive into the research topics of 'Latent 3D Scene Graph with Aligned Visual-Language Perception for Object-Goal Navigation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver