TY - JOUR
T1 - MLCA: Multi-level Correlative Attacks against Deep Cross-Modal Hashing
AU - Fang, Xiaohang
AU - Liu, Xin
AU - Hu, Zhikai
AU - Cheung, Yiu-ming
AU - Peng, Shu-Juan
AU - Xu, Xing
N1 - This work was supported in part by the National Science Foundation of China under Grants 62576143 and 62476103, in part by the National Science Foundation of Fujian Province under Grant 2024J01096, in part by the Xiamen Science and Technology Project under Grant 3502Z20251018, in part by the RGC Senior Research Fellow Scheme with the Grant SRFS2324-2S02, in part by the General Research Fund of RGC under the Grant 12202924, and in part by the RGC Junior Research Fellow Scheme under Grant JRFS2526-2S06.
Publisher Copyright:
© 2026 Elsevier Ltd
PY - 2026/6/1
Y1 - 2026/6/1
N2 - Deep Cross-Modal Hashing (DCMH) has gained significant popularity for its effectiveness in cross-modal retrieval. However, its intrinsic vulnerability to adversarial attacks poses a serious threat to retrieval reliability, highlighting the need for more comprehensive investigations of such attacks to improve model robustness. Despite recent progress, existing attack methods face two key limitations: (1) insufficiently exploit cross-modal adversarial correlations; (2) rely on indirect perturbation generation, which limits their ability to balance attack effectiveness and imperceptibility. To address these challenges, we propose an efficient Multi-level Correlative Attack (MLCA) against DCMH. Specifically, MLCA leverages the semantic extraction network and hash auxiliary network to extract multi-level information, from both clean and adversarial data. Guided by such cross-modal information, an efficient iterative attacking mechanism is innovatively designed to circularly manipulate cross-modal correlations of adversarial data, which seamlessly decouples the potential correlations within clean data while implanting deceptive cross-modal correlations into adversarial samples in Hamming space. Through the iterative refinement of the above process, adversarial examples with imperceptible perturbations can be effectively generated to attack various DCMH models. We evaluate MLCA on commonly used cross-modal retrieval benchmarks and representative DCMH models under the white-box setting. Experimental results show that MLCA generally achieves stronger targeted attack performance compared with baselines while maintaining low perceptual distortion. Furthermore, our robustness analysis reveals intrinsic vulnerabilities in existing models and provides valuable insights for developing robust cross-modal hashing systems. The code is available at: https://github.com/FunXH/MLCA.
AB - Deep Cross-Modal Hashing (DCMH) has gained significant popularity for its effectiveness in cross-modal retrieval. However, its intrinsic vulnerability to adversarial attacks poses a serious threat to retrieval reliability, highlighting the need for more comprehensive investigations of such attacks to improve model robustness. Despite recent progress, existing attack methods face two key limitations: (1) insufficiently exploit cross-modal adversarial correlations; (2) rely on indirect perturbation generation, which limits their ability to balance attack effectiveness and imperceptibility. To address these challenges, we propose an efficient Multi-level Correlative Attack (MLCA) against DCMH. Specifically, MLCA leverages the semantic extraction network and hash auxiliary network to extract multi-level information, from both clean and adversarial data. Guided by such cross-modal information, an efficient iterative attacking mechanism is innovatively designed to circularly manipulate cross-modal correlations of adversarial data, which seamlessly decouples the potential correlations within clean data while implanting deceptive cross-modal correlations into adversarial samples in Hamming space. Through the iterative refinement of the above process, adversarial examples with imperceptible perturbations can be effectively generated to attack various DCMH models. We evaluate MLCA on commonly used cross-modal retrieval benchmarks and representative DCMH models under the white-box setting. Experimental results show that MLCA generally achieves stronger targeted attack performance compared with baselines while maintaining low perceptual distortion. Furthermore, our robustness analysis reveals intrinsic vulnerabilities in existing models and provides valuable insights for developing robust cross-modal hashing systems. The code is available at: https://github.com/FunXH/MLCA.
KW - Hash auxiliary network
KW - Imperceptible perturbation
KW - Iterative attacking mechanism
KW - Multi-level correlative attack
UR - https://www.scopus.com/pages/publications/105040955109
U2 - 10.1016/j.patcog.2026.114124
DO - 10.1016/j.patcog.2026.114124
M3 - Journal article
SN - 0031-3203
VL - 180, Part B
JO - Pattern Recognition
JF - Pattern Recognition
M1 - 114124
ER -