Skip to main navigation Skip to search Skip to main content

Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning

  • Xiuxiu Qi
  • , Yu Yang
  • , Jiannong Cao*
  • , Luyao Bai
  • , Chongshan Fan
  • , Chengtai Cao
  • , Hongpeng Wang*
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Language-Conditioned Manipulation (LCM) facilitates human-robot interaction via Behavioral Cloning (BC), which learns control policies from human demonstrations and serves as a cornerstone of embodied AI. Overcoming compounding errors in sequential action decisions remains a central challenge to improving BC performance. Existing approaches mitigate compounding errors through data augmentation, expressive representation, or temporal abstraction. However, they suffer from physical discontinuities and semantic-physical misalignment, leading to inaccurate action cloning and intermittent execution. In this paper, we present Continuous vision-language-action Co-Learning with Semantic-Physical Alignment (CCoL), a novel BC framework that ensures temporally consistent execution and fine-grained semantic grounding. It generates robust and smooth action execution trajectories through continuous co-learning across vision, language, and proprioceptive inputs (i.e., robot internal states). Meanwhile, we anchor language semantics to visuomotor representations by a bidirectional cross-attention to learn contextual information for action generation, successfully overcoming the problem of semantic-physical misalignment. Extensive experiments show that CCoL achieves an average 8.0% relative improvement across three simulation suites, with up to 19.2% relative gain in human-demonstrated bimanual insertion tasks. Real-world tests on a 7-DoF robot further confirm CCoL’s generalization under unseen and noisy object states. © 2026, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved.
Original languageEnglish
Title of host publicationProceedings of the 40th Annual AAAI Conference on Artificial Intelligence
EditorsSven Koenig, Chad Jenkins, Matthew E. Taylor
Place of PublicationWashington, DC
PublisherAAAI Press
Pages24900-24908
Number of pages9
ISBN (Print)9781577359067, 1577359062
DOIs
Publication statusPublished - 2026
Event40th Annual AAAI Conference on Artificial Intelligence (AAAI 2026) - Singapore EXPO, Singapore, Singapore
Duration: 20 Jan 202627 Jan 2026
https://aaai.org/conference/aaai/aaai-26/

Publication series

NameProceedings of the AAAI Conference on Artificial Intelligence
Number29
Volume40
ISSN (Print)2159-5399
ISSN (Electronic)2374-3468

Conference

Conference40th Annual AAAI Conference on Artificial Intelligence (AAAI 2026)
Abbreviated titleAAAI-26
PlaceSingapore
CitySingapore
Period20/01/2627/01/26
Internet address

Funding

This work was supported by National High Level Hospital Clinical Research Funding (Grant 2025-PUMCH-D-005), National Natural Science Foundation of China (Grant 62473213), Sustainable Development Science and Technology Special Project of Shenzhen (Grant KCXFZ20230731100900002), Shenzhen Science and Technology Program (Grant KQTD20210811090143060), and Beijing Tianjin Hebei Basic Research Cooperation Special Project (Grant 24JCZXJC00060). We also acknowledge support in part from HK RGC General Research Fund (No.: PolyU 15235424), “Research on Key Technologies for Systematic Artificial Intelligence Agents” project under China Mobile Innovation and Research Institute (No.: R24114H7), and Research Institute for Artificial Intelligence of Things, The Hong Kong Polytechnic University. This work was partially conducted at the Robotics and Embodied Intelligence (REI) Lab, The Education University of Hong Kong.

RGC Funding Information

  • RGC-funded

Fingerprint

Dive into the research topics of 'Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning'. Together they form a unique fingerprint.

Cite this