Skip to main navigation Skip to search Skip to main content

Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation

  • Niantong Li
  • , Guangzheng Hu
  • , Weixu Qiao
  • , Ying Ba
  • , Qichen Hong
  • , Shijun Shen
  • , Jinlin Wang
  • , Fan Zhou
  • , Jianye Kang
  • , Xin Shang
  • , Ziyi He
  • , Wei Wang
  • , Dalin Li
  • , Jiahao Li
  • , Jie Zhang
  • , Kaiyuan Gao
  • , Kun Yan
  • , Lihan Jiang
  • , Ningyuan Tang
  • , Shengming Yin
  • Tianhe Wu, Xiao Xu, Xiaoyue Chen, Yuxiang Chen, Yan Shu, Yanran Zhang, Yilei Chen, Yixian Xu, Zekai Zhang, Zhendong Wang, Zihao Liu, Zikai Zhou, Hongzhu Shi, Yi Wang, Bing Zhao, Hu Wei, Lin Qu, Chenfei Wu*
*Corresponding author for this work

Research output: Working PapersWorking paper

Abstract

Text-to-Image (T2I) generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple text-image alignment can no longer satisfy users’ pressing demands for faithful real-world reconstruction and genuine creative expression. Existing benchmarks, however, remain anchored in these foundational criteria and do not yet capture the nuanced capabilities that matter in authentic artistic practice, making it difficult to reliably distinguish state-of-the-art T2I models. Moreover, many recent evaluation pipelines rely heavily on a unsupervised multimodal large language model (MLLM) as the sole judge, diverging from professional human standards. To address these gaps, we introduce Qwen-Image-Bench, a creator-centric benchmark co-designed with professional artists and grounded in real-world creation scenarios. Building upon the conventional pillars of Quality, Aesthetics, andText-Image Alignment, our benchmark enriches evaluation with two application-driven dimensions: Real-world Fidelity and Creative Generation. Drawing on the staged reasoning inherent in professional artistic workflows, we organize these five pillars into a top-down hierarchical taxonomy that further decomposes into 23 second-level sub-capabilities and 56 third-level verifiable rubrics. To ensure broad coverage, we curate 1,000 stratified bilingual prompts balanced across length and language, with each prompt jointly exercising more than four fine-grained facets across multiple pillars. We train a unified judge model (Q-Judger) based on Qwen3.6-27B, supervised by 80 professional annotators from art academies under blind labeling and triple-review protocols, that scores every image across all 56 verifiable facets, producing fine-grained, rubric-grounded, and fully attributable diagnostics rather than a single opaque score. Empirically, Qwen-Image-Bench reliably distinguishes leading T2I models, achieving the greatest separation on the two application-driven dimensions of Real-world Fidelity and Creative Generation where existing benchmarks provide little insight, while also providing a trustworthy optimization signal for production-level T2I development.
Original languageEnglish
Number of pages31
DOIs
Publication statusPublished - Jun 2026

Bibliographical note

Research Unit(s) information for this publication is provided by the author(s) concerned.

Fingerprint

Dive into the research topics of 'Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation'. Together they form a unique fingerprint.

Cite this