Skip to main navigation Skip to search Skip to main content

HumanEval-V: Systematic Evaluation of Visual Reasoning in Large Multimodal Models for Code Generation

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

Abstract

Visual representations are essential for communicating complex programming concepts and play a critical role in algorithm and code design. While recent advances in Large Multimodal Models (LMMs) have demonstrated impressive visual capabilities across various domains, their visual reasoning abilities in coding contexts remain largely unexplored. In this work, we present a systematic evaluation of LMMs’ visual reasoning abilities in coding contexts. We introduce HumanEval-V, a benchmark comprising 253 human-annotated code generation tasks, each requiring the generation of correct code solutions based on problem contexts encoded in diagrams. To ensure high-quality annotations, the construction of these tasks involves over 800 hours of meticulous human effort. Our tasks span six diverse categories and assess a broad spectrum of visual reasoning skills relevant to real-world programming scenarios. Through evaluation of 27 state-of-the-art LMMs, we find that even top-performing models such as Claude 3.5 Sonnet and Pixtral 124B achieve only 36.8% and 21.3% pass@1, respectively, while many open-weight models perform below 10%. Error analysis reveals that LMMs frequently generate hallucinations by misinterpreting or inventing visual details, and exhibit particular limitations in spatial reasoning, topological understanding, and handling dynamic visual patterns that are straightforward for humans. We finally discuss research opportunities and challenges to enhance model capabilities in visual reasoning for code generation. © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM.
Original languageEnglish
JournalACM Transactions on Software Engineering and Methodology
Online published15 May 2026
DOIs
Publication statusOnline published - 15 May 2026

Funding

This research was supported by National Natural Science Foundation of China under Grant No. 62502440 and Zhejiang Provincial Natural Science Foundation of China under Grant No. LQN26F020003.

Research Keywords

  • Large Language Model
  • Large Multimodal models
  • Visual Reasoning
  • Code Generation

Fingerprint

Dive into the research topics of 'HumanEval-V: Systematic Evaluation of Visual Reasoning in Large Multimodal Models for Code Generation'. Together they form a unique fingerprint.

Cite this