Skip to main navigation Skip to search Skip to main content

Visual Content Quality Assessment, Restoration, and Compression under Complex Degradation

Student thesis: Doctoral Thesis

Abstract

With the proliferation of multimedia technologies, the acquisition, transmission, and computation of visual content have become increasingly pervasive. However, in complex degradation scenarios, such as extreme low-light environments, historical analog films with composite damage, and extreme low-bandwidth communication channels, which cause visual signals inevitably suffer from severe quality degradation, structural corruption, and semantic loss. Traditional visual processing pipelines struggle to handle these extreme conditions. To tackle these challenges, this thesis systematically investigates the lifecycle of degraded visual content, establishing a comprehensive framework encompassing Visual Quality Assessment, Content Restoration, and Semantic-Perceptual Coding. By leveraging advanced generative priors and foundation models, this thesis bridges the gap between signal fidelity and high-level semantic perception. The main contributions are summarized into four interconnected works:

First, to establish a reliable perceptual metric for degraded and enhanced visual content, we investigate visual quality assessment in low-light enhancement scenarios. Existing quality metrics lack fine-grained perception for low-level features. To address this, we construct a large-scale Benchmark Dataset (RSLE-IQA) and propose a novel quality assessment framework based on contrastive prompt learning. By harnessing the robust semantic priors of Vision-Language Pre-training Models (VLMs), this work significantly improves the perceptual capacity of deep networks for low-level quality features, serving as the perceptual foundation for subsequent restoration and coding tasks.

Second, building upon the understanding of visual perception, we advance from assessment to active restoration, specifically targeting the highly complex composite degradations found in old films. Unlike modern digital videos, old films suffer from unpredictable analog and digital corruptions. We propose MambaOFR, a novel State Space Model (Mamba) based restoration framework. By dynamically generating degradation-aware prompts to adjust removal patterns and introducing a flow-guided mask deformable alignment module, this method effectively eradicates complex structural defects and prevents their temporal propagation. This work successfully reconstructs high-quality semantic content from severely physically degraded sources.

Third, as restored or captured visual content must often be transmitted under stringent channel constraints, we explore semantic-perceptual coding under ultra-low bitrates. Extreme bandwidth limitation acts as a severe "channel degradation," where traditional codecs fail to maintain both signal fidelity and machine-recognizable semantics. We propose a joint human and machine perception face compression framework that transmits only sketches and thumbnails. Utilizing a two-stage generative reconstruction mechanism with semi-parametric modeling and external database retrieval, this approach successfully preserves high human visual quality, identity consistency, and machine analysis performance at ultra-low bitrates.

Finally, to unify the ultimate goals of restoration (signal fidelity) and assessment (human perception), we propose a generalized, controllable coding framework leveraging Large Vision-Language Model (LVLM) priors. In neural image compression, striking a flexible balance between strict signal fidelity and subjective perceptual quality remains a critical challenge. We design a plug-and-play decoding framework featuring a scalable Low-Rank Adaptation (LoRA) scheme and a two-stage agent-assisted decoding strategy. By explicitly extracting and injecting textual semantic priors from LVLMs, this framework effectively eliminates semantic uncertainty during decoding, allowing existing neural codecs to controllably optimize the fidelity-perception trade-off.

In conclusion, this thesis provides a systematic and progressive solution for handling visual content in complex degradation scenarios. From establishing VLM-based quality assessment metrics to restoring composite physical degradations, and finally achieving generalized semantic-perceptual compression, the proposed methodologies offer profound theoretical insights and broad practical applications for next-generation resilient visual communication and computation.
Date of Award8 May 2026
Original languageEnglish
Awarding Institution
  • City University of Hong Kong
SupervisorShiqi WANG (Supervisor)

Keywords

  • Image compression
  • Image processing

Cite this

'