Abstract
Artistic Image Aesthetics Assessment (AIAA) aims to emulate human artistic perception to evaluate the aesthetics of artistic images. Due to the highly specialized nature of human artistic perception, obtaining large-scale aesthetic annotations for model analysis presents significant challenges. Furthermore, the subjectivity of artistic aesthetics makes it difficult for existing AIAA methods to quantify aesthetic scores solely based on visual features. To address the two challenges, we introduce an AIAA model based on multi-level text prompts generation. Firstly, we leverage a text prompt-based self-supervised learning approach to augment artistic image data and adopt a multi-task learning paradigm to pre-train our multi-modal AIAA model. To further capture the abstract aesthetic characteristics of artistic images, we then adopt a domain-specific multi-modal large language model (MLLM) to simulate human artistic perception and generate aesthetic textual descriptions for images, and employ a multi-modal fusion module to integrate image features with the text features of artistic aesthetics for better feature representation. Finally, the proposed multi-modal AIAA model is trained by text-prompt learning based on the aesthetic quality levels of artistic images. By generating multi-level text prompts, our method can introduce human-perceived artistic aesthetic knowledge and obtain a more effective AIAA model. Experimental results on several AIAA datasets demonstrate that our method is superior to the state-of-the-art AIAA methods.
© 2026 IEEE. All rights reserved, including rights for text and data mining and training of artificial intelligence and similar technologies. Personal use is permitted, but republication/redistribution requires IEEE permission.
© 2026 IEEE. All rights reserved, including rights for text and data mining and training of artificial intelligence and similar technologies. Personal use is permitted, but republication/redistribution requires IEEE permission.
| Original language | English |
|---|---|
| Number of pages | 12 |
| Journal | IEEE Transactions on Multimedia |
| DOIs | |
| Publication status | Online published - 18 May 2026 |
Funding
This work was supported in part by the National Natural Science Foundation of China under Grants 62101555, 62172417, 62471349, and 62171340, and in part by the Natural Science Foundation of Jiangsu Province under Grant BK20210488.
Research Keywords
- Artistic image aesthetics assessment
- artistic perception
- multi-level text prompts generation
- multi-modal learning
Fingerprint
Dive into the research topics of 'Learning Artistic Image Aesthetics from Multi-level Text Prompts Generation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver