Abstract
Visual representations are essential for communicating complex programming concepts and play a critical role in algorithm and code design. While recent advances in Large Multimodal Models (LMMs) have demonstrated impressive visual capabilities across various domains, their visual reasoning abilities in coding contexts remain largely unexplored. In this work, we present a systematic evaluation of LMMs’ visual reasoning abilities in coding contexts. We introduce HumanEval-V, a benchmark comprising 253 human-annotated code generation tasks, each requiring the generation of correct code solutions based on problem contexts encoded in diagrams. To ensure high-quality annotations, the construction of these tasks involves over 800 hours of meticulous human effort. Our tasks span six diverse categories and assess a broad spectrum of visual reasoning skills relevant to real-world programming scenarios. Through evaluation of 27 state-of-the-art LMMs, we find that even top-performing models such as Claude 3.5 Sonnet and Pixtral 124B achieve only 36.8% and 21.3% pass@1, respectively, while many open-weight models perform below 10%. Error analysis reveals that LMMs frequently generate hallucinations by misinterpreting or inventing visual details, and exhibit particular limitations in spatial reasoning, topological understanding, and handling dynamic visual patterns that are straightforward for humans. We finally discuss research opportunities and challenges to enhance model capabilities in visual reasoning for code generation. © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM.
| Original language | English |
|---|---|
| Journal | ACM Transactions on Software Engineering and Methodology |
| Online published | 15 May 2026 |
| DOIs | |
| Publication status | Online published - 15 May 2026 |
Funding
This research was supported by National Natural Science Foundation of China under Grant No. 62502440 and Zhejiang Provincial Natural Science Foundation of China under Grant No. LQN26F020003.
Research Keywords
- Large Language Model
- Large Multimodal models
- Visual Reasoning
- Code Generation
Fingerprint
Dive into the research topics of 'HumanEval-V: Systematic Evaluation of Visual Reasoning in Large Multimodal Models for Code Generation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver