Skip to main navigation Skip to search Skip to main content

面向具身智能的多模态视觉数据压缩综述

Translated title of the contribution: Overview of Multi-Modal Visual Data Compression for Embodied Intelligence
  • 王奕桐
  • , 张秋丹
  • , 张云
  • , 胡瑞珍
  • , 侯军辉
  • , 王旭*
  • *Corresponding author for this work

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

Abstract

Embodied artificial intelligence (AI) represents a paradigm shift from passive analysis to active interaction, serving as a critical pathway toward general AI. The emergence of the Vision-Language-Action (VLA) model has endowed robotic agents with unprecedented semantic understanding and generalization capabilities. However, the immense computational demands of these models necessitate an “Edge-Cloud Synergy” deployment architecture. This creates a major bottleneck: the massive data throughput generated by multimodal visual sensors far exceeds the limited and fluctuating bandwidth available in real-world scenarios. Traditional compression standards, designed for human-vision-oriented or machine-vision-oriented purposes, suffer from a fundamental objective mismatch in embodied contexts. Research indicates that artifacts introduced by these methods at low bitrates result in a significant performance degradation in task success rates and lead to dangerous error accumulation within the closed-loop control system. To address these urgent needs,this study provides a comprehensive survey of visual data compression and communication technologies tailored for embodied perception. First, we introduce the new paradigm of embodied-AI-oriented compression. We distinguish it from traditional paradigms by analyzing unique characteristics such as closed-loop error propagation, nonlinear performance responses, and the heterogeneous perceptual preference of downstream VLA models. Second, we systematically review state-of-the-art compression algorithms for planar vision, panoramic and fisheye vision, stereo vision, RGB-D data, and 3D point clouds, classified based on sensory modality. The survey places particular emphasis on Deep Neural Network-based innovations, including nonlinear transformation, end-to-end optimization, and cross-modal fusion, analyzing their adaptability to robotic manipulation and navigation tasks. Finally, we summarize existing challenges regarding latency and robustness and outline future research directions,including standardized evaluation benchmarks for embodied visual compression, task-driven dynamic compression, unified multimodal representation, end-to-end joint optimization of compression and control policies, and generalized compression for heterogeneous embodied models. © 2026, Editorial Board of Journal of Signal Processing. All rights reserved.
Translated title of the contributionOverview of Multi-Modal Visual Data Compression for Embodied Intelligence
Original languageChinese (Simplified)
Pages (from-to)1011-1028
Number of pages18
Journal信号处理
Volume42
Issue number7
DOIs
Publication statusPublished - Jul 2026

Funding

The National Natural Science Foundation of China (62371310); Shenzhen Science and Technology Program (JCYJ20241202124415021).

Research Keywords

  • deep neural network
  • embodied AI
  • vision-language-action models
  • visual data compression
  • 具身智能
  • 视觉数据压缩
  • 深度神经网络
  • 视觉-语言-动作模型

Fingerprint

Dive into the research topics of 'Overview of Multi-Modal Visual Data Compression for Embodied Intelligence'. Together they form a unique fingerprint.

Cite this