Skip to main navigation Skip to search Skip to main content

Transferability of Adversarial Attacks in Video-based MLLMs: A Cross-modal Image-to-Video Approach

  • Linhao Huang
  • , Xue Jiang
  • , Zhiqiang Wang
  • , Wentao Mo
  • , Xi Xiao*
  • , Yong Jie Yin
  • , Bo Han
  • , Feng Zheng*
  • *Corresponding author for this work

Research output: Chapter in book/report/conference proceedingConference proceedingpeer-review

Abstract

Video-based multimodal large language models (V-MLLMs) have shown vulnerability to adversarial examples in video-text multimodal tasks. However, the transferability of adversarial videos to unseen models—a common and practical real-world scenario—remains unexplored. In this paper, we pioneer an investigation into the transferability of adversarial video samples across V-MLLMs. We find that existing adversarial attack methods face significant limitations when applied in black-box settings for V-MLLMs, which we attribute to the following shortcomings: (1) lacking generalization in perturbing video features, (2) focusing only on sparse key-frames, and (3) failing to integrate multimodal information. To address these limitations and deepen the understanding of V-MLLM vulnerabilities in black-box scenarios, we introduce the Image-to-Video MLLM (I2V-MLLM) attack. In I2V-MLLM, we utilize an image-based multimodal large language model (I-MLLM) as a surrogate model to craft adversarial video samples. Multimodal interactions and spatiotemporal information are integrated to disrupt video representations within the latent space, improving adversarial transferability. Additionally, a perturbation propagation technique is introduced to handle different unknown frame sampling strategies. Experimental results demonstrate that our method can generate adversarial examples that exhibit strong transferability across different V-MLLMs on multiple video-text multimodal tasks. Compared to white-box attacks on these models, our black-box attacks (using BLIP-2 as a surrogate model) achieve competitive performance, with average attack success rate (AASR) of 57.98% on MSVD-QA and 58.26% on MSRVTT-QA for Zero-Shot VideoQA tasks, respectively.

Original languageEnglish
Title of host publicationProceedings of the 40th AAAI Conference on Artificial Intelligence, AAAI 2026
PublisherAAAI press
Pages5067-5075
Number of pages9
ISBN (Print)1577359062, 9781577359067
DOIs
Publication statusPublished - 17 Mar 2026
Event40th AAAI Conference on Artificial Intelligence, AAAI 2026 - Singapore, Singapore
Duration: 20 Jan 202627 Jan 2026
https://aaai.org/conference/aaai/aaai-26/ (Conference website)
https://aaai.org/conference/aaai/aaai-26/program-overview/ (Conference programme)
https://ojs.aaai.org/index.php/AAAI/index (Conference Proceedings )

Publication series

NameProceedings of the AAAI Conference on Artificial Intelligence
PublisherAAAI Press
Number7
Volume40
ISSN (Print)2159-5399

Conference

Conference40th AAAI Conference on Artificial Intelligence, AAAI 2026
Abbreviated titleAAAI 2026
Country/TerritorySingapore
CitySingapore
Period20/01/2627/01/26
Internet address

Fingerprint

Dive into the research topics of 'Transferability of Adversarial Attacks in Video-based MLLMs: A Cross-modal Image-to-Video Approach'. Together they form a unique fingerprint.

Cite this