Skip to main navigation Skip to search Skip to main content

Seeing is Believing: A Provable Path to Uncertainty Estimation in Large Vision-Language Models for Reliable Inference

Project: Research

Project Details

Description

Large Vision-Language Models (LVLMs) have recently emerged as powerful tools for a wide range of tasks, including visual question answering, multimodal reasoning, and image captioning. By coupling high-capacity vision encoders with large language models,LVLMs such as LLaVa demonstrate impressive capabilities in high-impact domains such as education, healthcare, and scientific research, where reliability and trust are critical. However, while LVLMs generate fluent and contextually rich responses, their expressions of uncertainty are always miscalibrated, exhibiting overconfidence even when predictions are incorrect. Such unreliable answers undermine user trust and limit the safe deployment of these models in real-world decision-making scenarios.Despite significant progress in model calibration that aligns predicted confidence with true accuracy in traditional classification tasks, a provable path to well-calibrated LVLMs remains an open problem. Unlike traditional closed-set classification with a single modality, LVLM tasks are inherently open-set and multimodal, where visual evidence and textual reasoning interact in complex ways. Existing calibration techniques follow standard training pipelines that rely on deterministic image-text pairs, implicitly teaching models that the world is certain, which introduces an intrinsic posterior error in confidence estimation. Furthermore, current methods cannot generalize calibration across diverse multimodal downstream tasks.In this project, we emphasize that seeing is believing, i.e., every textual claim generated by LVLMs should be grounded in reliable visual evidence. Based on this foundation, we investigate a roadmap towards reliable LVLMs, which contains three major tasks.We will theoretically study the posterior error in existing training pipelines and explicitly control it for provable calibration. By inverting the difficult posterior estimation problem, a new task of generating images conditioned on posterior probabilities can be introduced, which theoretically reduces calibration error with low sample complexity. We will leverage the synthetic data to pretrain a well-calibrated vision encoder in LVLMs. After parsing LVLM answers into unified claims, the uncertainty of each claim can be accurately quantified via reliable visual features from this vision encoder.We will integrate calibrated confidence into LVLMs for reliable inference, including visual reasoning that performs high-resolution reinspection for low-confidence regions and uncertainty expression with user interaction via confidence-guided clarification.We fulfill the research gap by investigating a new roadmap towards reliable LVLM instead of traditional calibration within limited scopes. A new data-driven paradigm for calibration research is introduced with theoretical guidance, which has the potential to accelerate the advancement of this research field.
Project number9048359
Grant typeECS
StatusNot started
Effective start/end date1/09/26 → …

Fingerprint

Explore the research topics touched on by this project. These labels are generated based on the underlying awards/grants. Together they form a unique fingerprint.