Abstract
Cross-view geo-localization aims to match images captured from different views over the same geographic region. Existing methods typically determine spatial correlations between cross-view images according to the similarity of representations extracted from salient areas. However, such local appearance representations fail to capture the underlying structural relationships among the corresponding regions, which severely undermines the reliability of localization results in complex scenes. To address this problem, we propose a reversible visual state-space model to enhance the understanding of global structural relations inherent in images captured from different views. Specifically, we design a progressive spatial analysis approach, which incrementally integrates geometric dependencies exploited at different levels to improve the understanding of the global structure. Moreover, we introduce a reversible rotational scanning mechanism based on the 2D-selective-scan (SS2D) module to facilitate the exploitation of geometric dependencies between cross-view images. Finally, we adopt the cross-dimension interaction strategy to enrich the informativeness of representations in the common space, thereby reinforcing the discriminability of cross-view representations between different regions. Extensive experiments on the University-1652 and University160k-WX datasets demonstrate that the proposed method achieves state-of-the-art performance while maintaining robustness under complex environmental conditions.
| Original language | English |
|---|---|
| Title of host publication | UAVM 2025 - Proceedings of the 3rd International Workshop on UAVs in Multimedia |
| Subtitle of host publication | Capturing the World from a New Perspective, Co-located with MM 2025 |
| Place of Publication | New York, NY, USA |
| Publisher | Association for Computing Machinery (ACM) |
| Pages | 42–46 |
| Number of pages | 5 |
| ISBN (Electronic) | 9798400718397 |
| DOIs | |
| Publication status | Published - 31 Oct 2025 |
Publication series
| Name | MM: International Multimedia Conference |
|---|---|
| Publisher | Association for Computing Machinery |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 9 Industry, Innovation, and Infrastructure
User-Defined Keywords
- Cross-view Geo-localization
- Structure Relation
- State Space Model
Fingerprint
Dive into the research topics of 'Understanding Global Structure Relation via Reversible Visual State Space Model for Robust Cross-View Geo-Localization'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver