Skip to main navigation Skip to search Skip to main content

CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Multimodal Large Language Models (MLLMs) achieve strong reasoning and perception capabilities but are increasingly vulnerable to jailbreak attacks. While existing work focuses on explicit attacks, where malicious content resides in a single modality, recent studies reveal implicit attacks, in which benign text and image inputs jointly express unsafe intent. Such joint-modal threats are difficult to detect and remain underexplored, largely due to the scarcity of high-quality implicit data. We propose ImpForge, an automated red-teaming pipeline that leverages reinforcement learning with tailored reward modules to generate diverse implicit samples across 14 domains. Building on this dataset, we further develop CrossGuard, an intent-aware safeguard providing robust and comprehensive defense against both explicit and implicit threats. Extensive experiments across safe and unsafe benchmarks, implicit and explicit attacks, and multiple out-of-domain settings demonstrate that CrossGuard significantly outperforms existing defenses, including advanced MLLMs and guardrails, achieving stronger security while maintaining high utility. This offers a balanced and practical solution for enhancing MLLM robustness against real-world multimodal threats.
Original languageEnglish
Title of host publicationProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Editors Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Place of PublicationSan Diego, California
PublisherAssociation for Computational Linguistics
Pages25693-25707
Number of pages15
Volume1
ISBN (Electronic)979-8-89176-390-6
DOIs
Publication statusPublished - Jul 2026
Event64th Annual Meeting of the Association for Computational Linguistics (ACL 2026) - San Diego, United States
Duration: 2 Jul 20267 Jul 2026
https://2026.aclweb.org

Conference

Conference64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)
Abbreviated titleACL 2026
PlaceUnited States
CitySan Diego
Period2/07/267/07/26
Internet address

Fingerprint

Dive into the research topics of 'CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks'. Together they form a unique fingerprint.

Cite this