Abstract
This thesis investigates crowd analysis from a frequency-domain perspective, with the goal of addressing structural limitations inherent in conventional spatial-domain approaches. While most existing crowd analysis methods operate directly on spatial representations, such formulations often suffer from loosely organized information structures, excessive degrees of freedom, and suboptimal utilization of supervisory signals. This thesis argues that these issues are not merely implementation-level challenges, but stem from the representational properties of the spatial domain itself.To overcome these limitations, we develop a principled theoretical framework that characterizes the transformation of spatial crowd representations into the frequency domain. We rigorously analyze the mathematical properties of frequency-domain representations, including information preservation, structural organization, stability under perturbations, and their interaction with supervision signals. Theoretical results are established to clarify under what conditions frequency-domain representations faithfully encode crowd density and localization information, and how these representations facilitate structured learning objectives. This foundation provides a formal justification for modeling crowd analysis tasks in the frequency domain rather than treating the transformation as a heuristic design choice.
Building upon this theoretical groundwork, we propose a unified frequency-domain framework for four representative crowd analysis tasks: crowd counting, crowd localization, noisy crowd counting, and video-based crowd counting. Instead of adapting spatial-domain loss functions directly, we design task-specific objective functions that explicitly leverage the structural characteristics of frequency-domain information. These objectives are derived and motivated by theoretical insights, ensuring consistency between representation properties and optimization strategies. For dynamic crowd scenarios, we further analyze temporal frequency structures and demonstrate how frequency-domain modeling enhances robustness and stability in video-based analysis.
Comprehensive theoretical analysis is complemented by extensive empirical validation on widely used benchmark datasets. Experimental results demonstrate that the proposed frequency-domain framework achieves competitive or superior performance compared to state-of-the-art spatial-domain approaches, while offering improved robustness to noise, enhanced supervision efficiency, and greater interpretability.
Overall, this thesis establishes frequency-domain modeling as a principled and theoretically grounded paradigm for crowd analysis. By unifying representation theory and task-specific learning within a coherent mathematical framework, it provides both conceptual insight and practical advances for future research in dense scene understanding and related computer vision problems.
| Date of Award | 15 May 2026 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Antoni Bert CHAN (Supervisor), Tak Wu Sam KWONG (Co-supervisor) & Kay Chen Tan (External Co-Supervisor) |
Cite this
- Standard