Abstract
Vision-language models (VLMs) have recently been adapted to image quality assessment (IQA), using language-defined quality levels to produce interpretable scores. This design introduces a new attack surface: readable text inside an image may be interpreted as evidence about quality rather than as ordinary scene content. However, the standard typographic attack paradigm, which pastes deceptive text directly onto an image, is poorly suited to IQA. A pasted overlay injects a positive semantic cue but simultaneously introduces a visible artifact that degrades image fidelity, creating a self-contradictory attack. We propose Naturalistic Typographic Attack (NTA), a black-box method that uses a text-guided image editor to embed quality-related words into scene-consistent carriers such as signs, posters, and labels. NTA preserves the semantic influence of typographic cues while avoiding the visual degradation of naive overlays. Experiments on AVA images scored by Q-Align show that NTA nearly doubles the mean score gain of direct overlay while achieving higher success rates and greater stability across images. These results indicate that scene-consistent text insertion exposes a more potent semantic vulnerability in VLM-based IQA than conventional pasted text.
| Original language | English |
|---|---|
| Title of host publication | 2026 18th International Conference on Quality of Multimedia Experience (QoMEX) |
| Publisher | IEEE |
| Pages | 1-4 |
| Number of pages | 4 |
| ISBN (Electronic) | 9798319500328 |
| ISBN (Print) | 9798319500335 |
| DOIs | |
| Publication status | Published - 29 Jun 2026 |
| Event | 2026 18th International Conference on Quality of Multimedia Experience (QoMEX) - Cardiff, United Kingdom Duration: 29 Jun 2026 → 3 Jul 2026 |
Publication series
| Name | International Workshop on Quality of Multimedia Experience, QoMEx |
|---|---|
| Publisher | IEEE |
| ISSN (Print) | 2372-7179 |
| ISSN (Electronic) | 2472-7814 |
Conference
| Conference | 2026 18th International Conference on Quality of Multimedia Experience (QoMEX) |
|---|---|
| Country/Territory | United Kingdom |
| City | Cardiff |
| Period | 29/06/26 → 3/07/26 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 9 Industry, Innovation, and Infrastructure
User-Defined Keywords
- adversarial robustness
- image quality assessment
- typographic attack
- vision-language model
Fingerprint
Dive into the research topics of 'Naturalistic Typographic Attacks on VLM-Based Image Quality Assessment'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver