Abstract
Red-teaming Vision-Language Models is essential for identifying vulnerabilities where adversarial image-text inputs trigger toxic outputs. Existing approaches treat image generation as a black box, providing only terminal toxicity scores while remaining temporally opaque regarding when and how toxic semantics emerge during multi-step synthesis. We introduce STARE, a hierarchical reinforcement learning framework that treats the denoising trajectory as an exploitable attack surface. By synergizing a high-level prompt editor with low-level T2I fine-tuning via Group Relative Policy Optimization (GRPO), STARE achieves a 68\% improvement in Attack Success Rate over state-of-the-art baselines including black box and white-box variants. More importantly, we reveal the Optimization-Induced Phase Alignment phenomenon: while vanilla models exhibit diffuse toxicity, adversarial optimization systematically concentrates conceptual harms into early semantic phases and detail-oriented harms into late refinement. This discovery transforms toxicity formation from a chaotic process into a series of predictable vulnerability windows. This temporal alignment transforms red-teaming from a trial-and-error process into a targeted structural analysis. Our work provides both a potent attack engine and a diagnostic foundation for developing next-generation, phase-aware safety mechanisms. Content warning: This paper contains examples of toxic content that may be offensive or disturbing.
© 2026 by the author(s)
© 2026 by the author(s)
| Original language | English |
|---|---|
| Title of host publication | Proceedings of the 43rd International Conference on Machine Learning (ICML 2026) |
| Publisher | ML Research Press |
| Publication status | Accepted/In press/Filed - 2026 |
| Event | 43rd International Conference on Machine Learning (ICML 2026) - Seoul, Korea, Republic of Duration: 6 Jul 2026 → 11 Jul 2026 https://icml.cc/Conferences/2026 |
Publication series
| Name | Proceedings of Machine Learning Research |
|---|---|
| ISSN (Electronic) | 2640-3498 |
Conference
| Conference | 43rd International Conference on Machine Learning (ICML 2026) |
|---|---|
| Abbreviated title | ICML 2026 |
| Place | Korea, Republic of |
| City | Seoul |
| Period | 6/07/26 → 11/07/26 |
| Internet address |
Bibliographical note
Since this conference is yet to commence, the information for this record is subject to revision. Full text of this publication does not contain sufficient affiliation information. With consent from the author(s) concerned, the Research Unit(s) information for this record is based on the existing academic department affiliation of the author(s).Fingerprint
Dive into the research topics of 'STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver