Abstract
With the proliferation of portable digital imaging devices like smartphones, images and videos have become the primary means of information perception. Current Global Shutter (GS) and Rolling Shutter (RS) sensors suffer from motion blur and RS distortion during motion. The emergence of Dynamic and Active-Pixel Vision Sensors (DAVIS) provides high dynamic range and temporal resolution event data, offering new spatiotemporal clues for quality restoration in both GS and RS videos. However, since most commercial cameras use RS, accelerating the adoption of novel sensors requires integrating event data with both GS and RS imaging devices. Existing event-driven deblurring algorithms for GS videos often overlook potential sharp textures and exposure time variations in real-world scenarios, limiting performance. Furthermore, current RS correction methods rely heavily on supervised training with synthetic data due to the difficulty of acquiring real paired data, which restricts their generalization ability. Addressing the above issues, this thesis focuses on the following main research contents:(1) To address the neglect of sharp textures in video deblurring, this thesis proposes a sharp-frame-guided framework. A bidirectional LSTM detector locates the Nearest Sharp Frame (NSF) to guide the deblurring process. A hybrid Transformer structure is introduced, combining local Transformers for adjacent frames and global Transformers for NSF integration. Additionally, a plug-and-play event fusion module merges event data with video features, bridging event-based and frame-based approaches. This enhances both performance and generalization.
(2) Existing event-based deblurring methods struggle with videos of varying exposure times. To overcome this, an exposure-aware deblurring and interpolation approach is proposed. It constructs an exposure-aware representation from blurry frames and event streams using supervised contrastive learning. Two U-Nets analyze intra- and inter-frame motion, adjusted by learned exposure features. The decoder uses exposure-adaptive convolution and gradual motion refinement to reconstruct videos across different blur levels. An event compression and filtering module separates exposure-related events, further aiding deblurring and interpolation. This method adapts to different exposures, significantly improving spatiotemporal quality.
(3) To overcome the reliance of existing methods on synthetic data and their poor generalization, this thesis proposes a self-supervised framework for dual reversed RS correction. The framework generates high-frame-rate GS videos directly from dual reversed RS images and their event streams. It filters event data, which is unaffected by RS distortion, to estimate the displacement between GS and RS images, guiding the correction network. A bidirectional distortion warping module enables self-supervised training by reconstructing the input RS images from corrected outputs, enforcing cycle consistency. A self-distillation loss is further incorporated to reduce boundary artifacts from warping. Together, these strategies improve correction quality and model generalization without paired supervision.
(4) To address the structural complexity and limited usability of current self-supervised RS correction methods, this thesis introduces a lightweight dual reversed RS correction network. Utilizing event guidance and a bidirectional correlation matching module, it jointly optimizes optical flow and correction features, enhancing performance while reducing parameters. A novel self-supervised strategy ensures cycle consistency between input and reconstructed RS images. Viewing RS reconstruction as a special case of video interpolation, each row of the RS image is interpolated from the corrected GS frame based on an RS distortion time map. This single-stage self-supervised approach achieves both high-quality RS correction and high-frame-rate GS video output.
By employing sharp texture search, event fusion, adaptive exposure perception, bidirectional warping modules, and self-supervised learning frameworks, this thesis effectively addresses key issues, including the lack of sharp textures in GS video deblurring, limitations in processing GS videos with varying exposure times, poor generalization of RS correction, and difficulties in acquiring paired data. These innovations significantly enhance the spatiotemporal quality of images and videos, providing robust support for the application of novel sensor technologies in practical scenarios.
| Date of Award | 16 Jan 2026 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Kede MA (Supervisor) & Dongwei Ren (External Supervisor) |
Cite this
- Standard