Abstract
Fine-Grained Sketch-Based Image Retrieval (FG-SBIR) aims to retrieve images that accurately correspond to abstract hand-drawn sketches, requiring the model to understand sparse and abstract visual cues. Existing methods tend to rely on convolutional networks or metric learning to align sketch and image features, often overlooking the inherent abstraction and semantic ambiguity present in sketches. This limitation results in an insufficient understanding of fine-grained visual details. To address this challenge, we propose SketchMind, a novel method that leverages Multi-modal Large Language Models (MLLMs) to enhance abstract sketch understanding in FG-SBIR. Specifically, we use MLLMs to generate auxiliary textual descriptions based on the given sketches via a Visual Question Answering (VQA) strategy. To effectively incorporate these descriptions, we construct a graph structure with the sketch as the central node and the generated texts as peripheral nodes. A graph attention scheme is employed to perform uncertainty-aware feature fusion, enabling the model to suppress noisy or irrelevant textual information. Furthermore, to enhance both inter- and intra-modal fine-grained alignment, we design a Multi-scale Cross-modal Jigsaw Matching module in combination with a self-supervised learning strategy, which captures local and global visual correspondences across modalities more effectively. Extensive experiments on three benchmark FG-SBIR datasets demonstrate that SketchMind achieves superior performance over existing state-of-the-art methods, proving its effectiveness. Code is available at https://github.com/li1changxing/MLLM_FG_SBIR/.
| Original language | English |
|---|---|
| Title of host publication | Proceedings of the ACM Web Conference, WWW 2026 |
| Place of Publication | New York, NY, USA |
| Publisher | Association for Computing Machinery (ACM) |
| Pages | 2453–2464 |
| Number of pages | 12 |
| ISBN (Electronic) | 9798400723070 |
| ISBN (Print) | 9798400723070 |
| DOIs | |
| Publication status | Published - 12 Apr 2026 |
| Event | 35th ACM Web Conference, WWW 2026 - Dubai, United Arab Emirates Duration: 13 Apr 2026 → 17 Apr 2026 https://dl.acm.org/doi/proceedings/10.1145/3774904 |
Publication series
| Name | Proceedings of the ACM Web Conference |
|---|---|
| Publisher | Association for Computing Machinery |
Conference
| Conference | 35th ACM Web Conference, WWW 2026 |
|---|---|
| Country/Territory | United Arab Emirates |
| City | Dubai |
| Period | 13/04/26 → 17/04/26 |
| Internet address |
User-Defined Keywords
- cross-modal retrieval
- hand-drawn sketch
- mllm
- MLLM
Fingerprint
Dive into the research topics of 'SketchMind: Understanding Abstract Sketches with MLLMs for Fine-Grained Sketch-Based Image Retrieval'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver