Skip to main navigation Skip to search Skip to main content

Learning Artistic Image Aesthetics from Multi-level Text Prompts Generation

  • Hancheng Zhu
  • , Xinya Xu
  • , Rui Yao
  • , Yixuan Li
  • , Kunyang Sun
  • , Leida Li

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

Abstract

Artistic Image Aesthetics Assessment (AIAA) aims to emulate human artistic perception to evaluate the aesthetics of artistic images. Due to the highly specialized nature of human artistic perception, obtaining large-scale aesthetic annotations for model analysis presents significant challenges. Furthermore, the subjectivity of artistic aesthetics makes it difficult for existing AIAA methods to quantify aesthetic scores solely based on visual features. To address the two challenges, we introduce an AIAA model based on multi-level text prompts generation. Firstly, we leverage a text prompt-based self-supervised learning approach to augment artistic image data and adopt a multi-task learning paradigm to pre-train our multi-modal AIAA model. To further capture the abstract aesthetic characteristics of artistic images, we then adopt a domain-specific multi-modal large language model (MLLM) to simulate human artistic perception and generate aesthetic textual descriptions for images, and employ a multi-modal fusion module to integrate image features with the text features of artistic aesthetics for better feature representation. Finally, the proposed multi-modal AIAA model is trained by text-prompt learning based on the aesthetic quality levels of artistic images. By generating multi-level text prompts, our method can introduce human-perceived artistic aesthetic knowledge and obtain a more effective AIAA model. Experimental results on several AIAA datasets demonstrate that our method is superior to the state-of-the-art AIAA methods.

© 2026 IEEE. All rights reserved, including rights for text and data mining and training of artificial intelligence and similar technologies. Personal use is permitted, but republication/redistribution requires IEEE permission.
Original languageEnglish
Number of pages12
JournalIEEE Transactions on Multimedia
DOIs
Publication statusOnline published - 18 May 2026

Funding

This work was supported in part by the National Natural Science Foundation of China under Grants 62101555, 62172417, 62471349, and 62171340, and in part by the Natural Science Foundation of Jiangsu Province under Grant BK20210488.

Research Keywords

  • Artistic image aesthetics assessment
  • artistic perception
  • multi-level text prompts generation
  • multi-modal learning

Fingerprint

Dive into the research topics of 'Learning Artistic Image Aesthetics from Multi-level Text Prompts Generation'. Together they form a unique fingerprint.

Cite this