Abstract
Image captioning, which converts visual input into natural language descriptions, is a fundamental cross-modal task with wide-ranging applications, including image understanding, automatic labeling, and human-computer interaction. While deep learning models have achieved considerable success in generating accurate image descriptions, they often fall short in producing distinctive captions--descriptions that uniquely capture the specific content of an image and differentiate it from similar images. This limitation poses a significant challenge in applications requiring fine-grained or large-scale image captioning, where uniquely describing each image is essential.(1) This dissertation addresses the challenge of enhancing the distinctiveness of image captions using novel generative models. Distinctiveness, as defined in this work, refers to the capacity of a caption to highlight the unique elements of a target image, distinguishing it from other images with similar content. To tackle this problem, we focus on generating distinct image captions within groups of visually similar images. We define image similarity and propose methods to group similar images effectively. Building on this, new distinctiveness evaluation metrics are introduced, which are specifically designed to measure how well a caption differentiates a target image from other images in the group.
(2) Key contributions of this work include the development of distinctive captioning methods based on generative models and innovative techniques for enhancing captioning distinctiveness through label weight adjustment and region-based attention mechanisms. The proposed label weight adjustment method addresses the problem of overfitting to high-frequency labels by introducing strategies to amplify the impact of distinctive labels. This method incorporates three key strategies: sentence weighting, long-tail word weighting, and negative sampling. Sentence weighting prioritizes discriminative sentences based on their distinctiveness score (CIDErBtw), while long-tail word weighting increases the importance of low-frequency words that are more likely to be informative. Negative sampling ensures that captions are dissimilar to those of similar images, encouraging more unique descriptions.
(3) Furthermore, this dissertation introduces a regional attention enhancement technique designed to increase focus on distinctive regions of an image during the encoding process. This method employs a Group-based Memory Attention (GMA) module, which compares memory vectors across similar images to highlight distinctive features in the target image. Additionally, loss functions are optimized to encourage the use of distinctive words during caption generation, thereby improving the overall distinctiveness of the generated captions. The experimental evaluation demonstrates the effectiveness of these methods. Three novel metrics (i.e., CIDErBtw, CIDErRank, and DisWordRate) are introduced to evaluate distinctiveness, and experimental results show that these metrics correlate strongly with human judgment. When applied to Transformer-based baseline models, the proposed techniques improve both caption accuracy and distinctiveness. Specifically, the label weight adjustment approach increases the baseline model's CIDEr score by 3.7 while reducing CIDErBtw by 5.2, indicating more distinctive captioning. Similarly, the regional attention enhancement technique improves the DisWordRate from 16.8 to 19.5, bringing it closer to human performance in generating unique captions.
(4) The recent advancements in generative models are also explored, focusing on diffusion models for multi-modal content generation. Image captioning is integrated as a subtask within the larger framework of the Omnidirectional Cross-modal Generative (OCG) model. The OCG model aligns visual, textual, and audio modalities through a diffusion-based approach that fuses multi-modal features to generate highly aligned content across different modalities. For image captioning, this cross-modal feature diffusion mechanism allows the generation of captions that are not only accurate but also distinctive by incorporating both visual and textual cues effectively. By leveraging diffusion models, the dissertation demonstrates that multi-modal generative models can produce high-quality, distinctive captions, even with limited training data. The OCG model excels in multi-modal tasks, including captioning and multi-modal pair generation, further enriching the scope of distinctive image captioning.
In summary, this dissertation makes significant advancements in the field of image captioning by addressing the long-standing issue of distinctiveness in generated captions. By leveraging generative models and proposing novel strategies for label weighting and regional attention, this work paves the way for more precise and distinctive image captioning, with potential applications in areas requiring fine-grained image understanding and multi-modal content generation.
| Date of Award | 24 Mar 2025 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Antoni Bert CHAN (Supervisor) & Yirong Wu (External Supervisor) |
Cite this
- Standard