Skip to main navigation Skip to search Skip to main content

STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack

  • Xutao Mao
  • , Liangjie Zhao
  • , Tao Liu
  • , Xiang Zheng*
  • , Hongying Zan
  • , Cong WANG*
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Red-teaming Vision-Language Models is essential for identifying vulnerabilities where adversarial image-text inputs trigger toxic outputs. Existing approaches treat image generation as a black box, providing only terminal toxicity scores while remaining temporally opaque regarding when and how toxic semantics emerge during multi-step synthesis. We introduce STARE, a hierarchical reinforcement learning framework that treats the denoising trajectory as an exploitable attack surface. By synergizing a high-level prompt editor with low-level T2I fine-tuning via Group Relative Policy Optimization (GRPO), STARE achieves a 68\% improvement in Attack Success Rate over state-of-the-art baselines including black box and white-box variants. More importantly, we reveal the Optimization-Induced Phase Alignment phenomenon: while vanilla models exhibit diffuse toxicity, adversarial optimization systematically concentrates conceptual harms into early semantic phases and detail-oriented harms into late refinement. This discovery transforms toxicity formation from a chaotic process into a series of predictable vulnerability windows. This temporal alignment transforms red-teaming from a trial-and-error process into a targeted structural analysis. Our work provides both a potent attack engine and a diagnostic foundation for developing next-generation, phase-aware safety mechanisms. Content warning: This paper contains examples of toxic content that may be offensive or disturbing.

© 2026 by the author(s)
Original languageEnglish
Title of host publicationProceedings of the 43rd International Conference on Machine Learning (ICML 2026)
PublisherML Research Press
Publication statusAccepted/In press/Filed - 2026
Event43rd International Conference on Machine Learning (ICML 2026) - Seoul, Korea, Republic of
Duration: 6 Jul 202611 Jul 2026
https://icml.cc/Conferences/2026

Publication series

NameProceedings of Machine Learning Research
ISSN (Electronic)2640-3498

Conference

Conference43rd International Conference on Machine Learning (ICML 2026)
Abbreviated titleICML 2026
PlaceKorea, Republic of
CitySeoul
Period6/07/2611/07/26
Internet address

Bibliographical note

Since this conference is yet to commence, the information for this record is subject to revision. Full text of this publication does not contain sufficient affiliation information. With consent from the author(s) concerned, the Research Unit(s) information for this record is based on the existing academic department affiliation of the author(s).

Fingerprint

Dive into the research topics of 'STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack'. Together they form a unique fingerprint.

Cite this