Abstract
Image Compression for Machines (ICM) aims to compress images for machine vision tasks, while current methods mostly focus on the demands for high-level tasks. However, the quality of original images is usually not guaranteed in the real world, leading to even worse downstream task performance after compression. Thus, lowlevel (LL) restoration tasks should also be considered in ICM. In this paper, we propose the first ICM framework for LL machine vision tasks, namely LL-ICM, which optimizes the compression and LL processing performance simultaneously. Moreover, LL-ICM leverages large vision-language model (VLM) to solve different LL task within a single model, which is particularly useful when the distortion type of the original image is uncertain. As illustrated in Fig. 1(a), LL-ICM consists of a neural image codec and a VLM-based LL processing module. Given an original image with distortions, LL-ICM firstly compress it as X^. Then, we extract a generalized feature F from X^, which is then encoded as two representations, distortion type φ and caption σ. After that, the LL processing module receives X^ and its representations to generate the restored version of X^, i.e., X^H. © 2025 IEEE
| Original language | English |
|---|---|
| Title of host publication | Proceedings - 2025 DCC |
| Subtitle of host publication | 2025 Data Compression Conference |
| Editors | Ali Bilgin, James E. Fowler |
| Publisher | IEEE |
| Pages | 408 |
| Number of pages | 1 |
| ISBN (Electronic) | 979-8-3315-3471-4 |
| ISBN (Print) | 979-8-3315-3472-1 |
| DOIs | |
| Publication status | Published - 2025 |
| Event | 2025 Data Compression Conference (DCC 2025) - Cliff Lodge convention center, Salt Lake City, United States Duration: 18 Mar 2025 → 21 Mar 2025 https://datacompressionconference.org/ |
Publication series
| Name | |
|---|---|
| ISSN (Print) | 1068-0314 |
| ISSN (Electronic) | 2375-0359 |
Conference
| Conference | 2025 Data Compression Conference (DCC 2025) |
|---|---|
| Abbreviated title | DCC 2025 |
| Place | United States |
| City | Salt Lake City |
| Period | 18/03/25 → 21/03/25 |
| Internet address |
Funding
This work was supported in part by the National Natural Science Foundation of China 62072008, and in part by the Shenzhen Science and Technology Program under Project JCYJ20220530140816037.
Fingerprint
Dive into the research topics of 'LL-ICM: Image Compression for Low-level Machine Vision via Large Vision-Language Model'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver