TY - UNPB
T1 - Qwen-Image-Bench
T2 - From Generation to Creation in Text-to-Image Evaluation
AU - Li, Niantong
AU - Hu, Guangzheng
AU - Qiao, Weixu
AU - Ba, Ying
AU - Hong, Qichen
AU - Shen, Shijun
AU - Wang, Jinlin
AU - Zhou, Fan
AU - Kang, Jianye
AU - Shang, Xin
AU - He, Ziyi
AU - Wang, Wei
AU - Li, Dalin
AU - Li, Jiahao
AU - Zhang, Jie
AU - Gao, Kaiyuan
AU - Yan, Kun
AU - Jiang, Lihan
AU - Tang, Ningyuan
AU - Yin, Shengming
AU - Wu, Tianhe
AU - Xu, Xiao
AU - Chen, Xiaoyue
AU - Chen, Yuxiang
AU - Shu, Yan
AU - Zhang, Yanran
AU - Chen, Yilei
AU - Xu, Yixian
AU - Zhang, Zekai
AU - Wang, Zhendong
AU - Liu, Zihao
AU - Zhou, Zikai
AU - Shi, Hongzhu
AU - Wang, Yi
AU - Zhao, Bing
AU - Wei, Hu
AU - Qu, Lin
AU - Wu, Chenfei
N1 - Research Unit(s) information for this publication is provided by the author(s) concerned.
PY - 2026/6
Y1 - 2026/6
N2 - Text-to-Image (T2I) generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple text-image alignment can no longer satisfy users’ pressing demands for faithful real-world reconstruction and genuine creative expression. Existing benchmarks, however, remain anchored in these foundational criteria and do not yet capture the nuanced capabilities that matter in authentic artistic practice, making it difficult to reliably distinguish state-of-the-art T2I models. Moreover, many recent evaluation pipelines rely heavily on a unsupervised multimodal large language model (MLLM) as the sole judge, diverging from professional human standards. To address these gaps, we introduce Qwen-Image-Bench, a creator-centric benchmark co-designed with professional artists and grounded in real-world creation scenarios. Building upon the conventional pillars of Quality, Aesthetics, andText-Image Alignment, our benchmark enriches evaluation with two application-driven dimensions: Real-world Fidelity and Creative Generation. Drawing on the staged reasoning inherent in professional artistic workflows, we organize these five pillars into a top-down hierarchical taxonomy that further decomposes into 23 second-level sub-capabilities and 56 third-level verifiable rubrics. To ensure broad coverage, we curate 1,000 stratified bilingual prompts balanced across length and language, with each prompt jointly exercising more than four fine-grained facets across multiple pillars. We train a unified judge model (Q-Judger) based on Qwen3.6-27B, supervised by 80 professional annotators from art academies under blind labeling and triple-review protocols, that scores every image across all 56 verifiable facets, producing fine-grained, rubric-grounded, and fully attributable diagnostics rather than a single opaque score. Empirically, Qwen-Image-Bench reliably distinguishes leading T2I models, achieving the greatest separation on the two application-driven dimensions of Real-world Fidelity and Creative Generation where existing benchmarks provide little insight, while also providing a trustworthy optimization signal for production-level T2I development.
AB - Text-to-Image (T2I) generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple text-image alignment can no longer satisfy users’ pressing demands for faithful real-world reconstruction and genuine creative expression. Existing benchmarks, however, remain anchored in these foundational criteria and do not yet capture the nuanced capabilities that matter in authentic artistic practice, making it difficult to reliably distinguish state-of-the-art T2I models. Moreover, many recent evaluation pipelines rely heavily on a unsupervised multimodal large language model (MLLM) as the sole judge, diverging from professional human standards. To address these gaps, we introduce Qwen-Image-Bench, a creator-centric benchmark co-designed with professional artists and grounded in real-world creation scenarios. Building upon the conventional pillars of Quality, Aesthetics, andText-Image Alignment, our benchmark enriches evaluation with two application-driven dimensions: Real-world Fidelity and Creative Generation. Drawing on the staged reasoning inherent in professional artistic workflows, we organize these five pillars into a top-down hierarchical taxonomy that further decomposes into 23 second-level sub-capabilities and 56 third-level verifiable rubrics. To ensure broad coverage, we curate 1,000 stratified bilingual prompts balanced across length and language, with each prompt jointly exercising more than four fine-grained facets across multiple pillars. We train a unified judge model (Q-Judger) based on Qwen3.6-27B, supervised by 80 professional annotators from art academies under blind labeling and triple-review protocols, that scores every image across all 56 verifiable facets, producing fine-grained, rubric-grounded, and fully attributable diagnostics rather than a single opaque score. Empirically, Qwen-Image-Bench reliably distinguishes leading T2I models, achieving the greatest separation on the two application-driven dimensions of Real-world Fidelity and Creative Generation where existing benchmarks provide little insight, while also providing a trustworthy optimization signal for production-level T2I development.
U2 - 10.48550/arXiv.2605.28091
DO - 10.48550/arXiv.2605.28091
M3 - Working paper
BT - Qwen-Image-Bench
ER -