A Parameter-Efficient Mixture-of-Experts Framework for Cross-Modal Geo-Localization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912665857687552 |
|---|---|
| author | Li, LinFeng Zhao, Jian Yang, Zepeng Song, Yuhang Lin, Bojun Zhang, Tianle Yuan, Yuchen Zhang, Chi Li, Xuelong |
| author_facet | Li, LinFeng Zhao, Jian Yang, Zepeng Song, Yuhang Lin, Bojun Zhang, Tianle Yuan, Yuchen Zhang, Chi Li, Xuelong |
| contents | We present a winning solution to RoboSense 2025 Track 4: Cross-Modal Drone Navigation. The task retrieves the most relevant geo-referenced image from a large multi-platform corpus (satellite/drone/ground) given a natural-language query. Two obstacles are severe inter-platform heterogeneity and a domain gap between generic training descriptions and platform-specific test queries. We mitigate these with a domain-aligned preprocessing pipeline and a Mixture-of-Experts (MoE) framework: (i) platform-wise partitioning, satellite augmentation, and removal of orientation words; (ii) an LLM-based caption refinement pipeline to align textual semantics with the distinct visual characteristics of each platform. Using BGE-M3 (text) and EVA-CLIP (image), we train three platform experts using a progressive two-stage, hard-negative mining strategy to enhance discriminative power, and fuse their scores at inference. The system tops the official leaderboard, demonstrating robust cross-modal geo-localization under heterogeneous viewpoints. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_20291 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Parameter-Efficient Mixture-of-Experts Framework for Cross-Modal Geo-Localization Li, LinFeng Zhao, Jian Yang, Zepeng Song, Yuhang Lin, Bojun Zhang, Tianle Yuan, Yuchen Zhang, Chi Li, Xuelong Computer Vision and Pattern Recognition Artificial Intelligence We present a winning solution to RoboSense 2025 Track 4: Cross-Modal Drone Navigation. The task retrieves the most relevant geo-referenced image from a large multi-platform corpus (satellite/drone/ground) given a natural-language query. Two obstacles are severe inter-platform heterogeneity and a domain gap between generic training descriptions and platform-specific test queries. We mitigate these with a domain-aligned preprocessing pipeline and a Mixture-of-Experts (MoE) framework: (i) platform-wise partitioning, satellite augmentation, and removal of orientation words; (ii) an LLM-based caption refinement pipeline to align textual semantics with the distinct visual characteristics of each platform. Using BGE-M3 (text) and EVA-CLIP (image), we train three platform experts using a progressive two-stage, hard-negative mining strategy to enhance discriminative power, and fuse their scores at inference. The system tops the official leaderboard, demonstrating robust cross-modal geo-localization under heterogeneous viewpoints. |
| title | A Parameter-Efficient Mixture-of-Experts Framework for Cross-Modal Geo-Localization |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2510.20291 |