Skip to main navigation Skip to search Skip to main content

LL-ICM: Image Compression for Low-level Machine Vision via Large Vision-Language Model

  • Yuan Xue
  • , Qi Zhang*
  • , Chuanmin Jia
  • , Shiqi Wang
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Image Compression for Machines (ICM) aims to compress images for machine vision tasks, while current methods mostly focus on the demands for high-level tasks. However, the quality of original images is usually not guaranteed in the real world, leading to even worse downstream task performance after compression. Thus, lowlevel (LL) restoration tasks should also be considered in ICM. In this paper, we propose the first ICM framework for LL machine vision tasks, namely LL-ICM, which optimizes the compression and LL processing performance simultaneously. Moreover, LL-ICM leverages large vision-language model (VLM) to solve different LL task within a single model, which is particularly useful when the distortion type of the original image is uncertain. As illustrated in Fig. 1(a), LL-ICM consists of a neural image codec and a VLM-based LL processing module. Given an original image with distortions, LL-ICM firstly compress it as X^. Then, we extract a generalized feature F from X^, which is then encoded as two representations, distortion type φ and caption σ. After that, the LL processing module receives X^ and its representations to generate the restored version of X^, i.e., X^H. © 2025 IEEE
Original languageEnglish
Title of host publicationProceedings - 2025 DCC
Subtitle of host publication2025 Data Compression Conference
EditorsAli Bilgin, James E. Fowler
PublisherIEEE
Pages408
Number of pages1
ISBN (Electronic)979-8-3315-3471-4
ISBN (Print)979-8-3315-3472-1
DOIs
Publication statusPublished - 2025
Event2025 Data Compression Conference (DCC 2025) - Cliff Lodge convention center, Salt Lake City, United States
Duration: 18 Mar 202521 Mar 2025
https://datacompressionconference.org/

Publication series

Name
ISSN (Print)1068-0314
ISSN (Electronic)2375-0359

Conference

Conference2025 Data Compression Conference (DCC 2025)
Abbreviated titleDCC 2025
PlaceUnited States
CitySalt Lake City
Period18/03/2521/03/25
Internet address

Funding

This work was supported in part by the National Natural Science Foundation of China 62072008, and in part by the Shenzhen Science and Technology Program under Project JCYJ20220530140816037.

Fingerprint

Dive into the research topics of 'LL-ICM: Image Compression for Low-level Machine Vision via Large Vision-Language Model'. Together they form a unique fingerprint.

Cite this