Abstract
With the rapid development of deep learning methods, object detection, which aims at localizing and categorizing objects with given different source data such as images or point clouds, has witnessed significant progress in different scenes in recent years, and has also gradually served as a powerful and reliable technology for many high-level applications, such as object tracking, surveillance systems, and autonomous driving. However, the impressive performance of object detection is at the huge cost of high-quality manual annotations, and will suffer a severe deterioration once the accurate annotations are insufficient, especially in complex scenes, such as the driving scene or aerial scene. Specifically, in the driving scene, the target objects are either small or non-static since the data are captured by a car-mounted camera. Limited annotations degrade the recognition performance of detection models on challenging moving and small objects. In addition, in the aerial scene, the orientation property is an essential property to characterize oriented objects in aerial images, and the detection methods should achieve accurate location and orientation estimation. The detection performance deteriorates when only limited annotations with precise angular targets are accessible. Semi-supervised learning is a promising method to address the problem of annotation deficiency when sufficient and low-cost unannotated data are available. Although semi-supervised learning has been extensively applied for object detection in the natural scene, it still encounters unusual challenges in both the driving scene and aerial scene due to the moving or orientation properties in different scenes. Therefore, in this thesis, we propose four methods to achieve semi-supervised learning in both the driving scene and the aerial scene. The main contributions of this thesis can be summarized as follows.1. In the first work, we propose to solve code-level information transferring from reliable domains to unreliable domains by incorporating a domain classifier that competes with the disentangling module to generate domain-invariant codes. An external classifier is trained on appearance-enhanced instances and sends completeness signals to the generative module, which facilitates the generative module to recognize truncated or non-truncated pedestrian instances. The resulting classifier ultimately generates high-quality pseudo-annotations for the unannotated data. The pseudo-annotated data, combined with a small amount of manually annotated data, are used to achieve a detector with improved generalization and accuracy. We perform extensive experiments on multiple challenging benchmarks to demonstrate the effectiveness of the proposed method.
2. In the second work, we propose to achieve semi-supervised learning for stereo-based 3D object detection through pseudo annotation generation from a temporal-aggregated teacher model, which temporally accumulates knowledge from a student model. To facilitate a more stable and accurate depth estimation, we introduce Temporal-Aggregation-Guided (TAG) disparity consistency, a cross-view disparity consistency constraint between the teacher model and the student model for robust and improved depth estimation. To mitigate noise in pseudo annotation generation, we propose a cross-view agreement strategy, in which pseudo annotations should attain high degree of agreements between 3D and 2D views, as well as between binocular views. We perform extensive experiments on the KITTI 3D dataset to demonstrate our proposed method’s capability in leveraging a huge amount of unannotated stereo images to attain significantly improved detection results.
3. In the third work, we propose Pseudo-Siamese Teacher (PST), a new semi-supervised learning framework for oriented object detection. In this architecture, two teacher models, updated from the same student model with different optimizations, inspect the predictions of each other and collaborate to generate high-quality pseudo annotations. To reduce the unreliability of pseudo annotations on the localization, scale, and orientation, we propose to model the oriented object as a Gaussian distribution and apply a symmetric and bounded Jensen–Shannon divergence (JSD) to evaluate the divergence between predictions of different teacher models, the results of which serve as an indicator to remove confusing pseudo annotations without consistent regression estimation of teacher models. Scale invariance is also an important challenge in oriented object detection, which we address by proposing a scale-adaptive knowledge distillation to align information between the feature maps from the student model on images with flexible scales and the feature maps, interpolated from adjacent feature maps with scales closest to those of the down-sampled images, from the teacher models. We perform extensive experiments to demonstrate the effectiveness of our proposed method in leveraging unannotated data for performance improvement.
4. In the fourth work, motivated by weakly supervised learning, we introduce annotation-efficient point annotations for unannotated images and propose a weakly semi-supervised method for oriented object detection to balance the detection performance and annotation cost. Specifically, we propose a Rotation-Modulated Relational Graph Matching method to match relations of proposals centered on annotated points between the teacher and student models to alleviate the ambiguity of point annotations in depicting the oriented object. In addition, we further propose a Relational Rank Distribution Matching method to align the rank distribution on classification and regression between different models. Finally, to handle the difficult annotated points that both models are confused about, we introduce weakly supervised learning to impose positive signals for difficult point-induced clusters to the base model, and focus the base model on the occupancy between the predictions and annotated points. We perform extensive experiments on challenging datasets to demonstrate the effectiveness of our proposed weakly semi-supervised method in leveraging point-annotated data for significant performance improvement.
| Date of Award | 28 Jan 2026 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Hau San WONG (Supervisor) |
Cite this
- Standard