REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yong, Jia, Furong, Yin, Dacheng, Rong, Kang, Rao, Fengyun, Lyu, Jing, Zhang, Fan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
by: Yang, Jie, et al.
Published: (2025)
by: Yang, Jie, et al.
Published: (2025)
SAIL: Self-Amplified Iterative Learning for Diffusion Model Alignment with Minimal Human Feedback
by: He, Xiaoxuan, et al.
Published: (2026)
by: He, Xiaoxuan, et al.
Published: (2026)
WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
by: Yang, Jian, et al.
Published: (2025)
by: Yang, Jian, et al.
Published: (2025)
D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning
by: Tang, Changli, et al.
Published: (2026)
by: Tang, Changli, et al.
Published: (2026)
Gen-Searcher: Reinforcing Agentic Search for Image Generation
by: Feng, Kaituo, et al.
Published: (2026)
by: Feng, Kaituo, et al.
Published: (2026)
TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
by: He, Xiaoxuan, et al.
Published: (2025)
by: He, Xiaoxuan, et al.
Published: (2025)
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
by: Yue, Xinli, et al.
Published: (2025)
by: Yue, Xinli, et al.
Published: (2025)
MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling
by: Yang, Jian, et al.
Published: (2024)
by: Yang, Jian, et al.
Published: (2024)
Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
by: Wang, Zitian, et al.
Published: (2025)
by: Wang, Zitian, et al.
Published: (2025)
ObjEmbed: Towards Universal Multimodal Object Embeddings
by: Fu, Shenghao, et al.
Published: (2026)
by: Fu, Shenghao, et al.
Published: (2026)
WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
Cross-view geo-localization, Image retrieval, Multiscale geometric modeling, Frequency domain enhancement
by: Zhang, Hongying, et al.
Published: (2026)
by: Zhang, Hongying, et al.
Published: (2026)
Towards Interpretable Geo-localization: a Concept-Aware Global Image-GPS Alignment Framework
by: Jia, Furong, et al.
Published: (2025)
by: Jia, Furong, et al.
Published: (2025)
Cross-view geo-localization: a survey
by: Durgam, Abhilash, et al.
Published: (2024)
by: Durgam, Abhilash, et al.
Published: (2024)
Semantic-Enriched Latent Visual Reasoning
by: Xu, Tianrun, et al.
Published: (2026)
by: Xu, Tianrun, et al.
Published: (2026)
R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction
by: Yan, Shannan, et al.
Published: (2026)
by: Yan, Shannan, et al.
Published: (2026)
SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning
by: Chng, Yong Xien, et al.
Published: (2025)
by: Chng, Yong Xien, et al.
Published: (2025)
Mind-Brush: Integrating Agentic Cognitive Search and Reasoning into Image Generation
by: He, Jun, et al.
Published: (2026)
by: He, Jun, et al.
Published: (2026)
Agentic Keyframe Search for Video Question Answering
by: Fan, Sunqi, et al.
Published: (2025)
by: Fan, Sunqi, et al.
Published: (2025)
Cross-view image geo-localization with Panorama-BEV Co-Retrieval Network
by: Ye, Junyan, et al.
Published: (2024)
by: Ye, Junyan, et al.
Published: (2024)
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
by: Wei, Zhixiang, et al.
Published: (2025)
by: Wei, Zhixiang, et al.
Published: (2025)
HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization
by: Zhou, Zitang, et al.
Published: (2025)
by: Zhou, Zitang, et al.
Published: (2025)
Spatial-Semantic Collaborative Cropping for User Generated Content
by: Su, Yukun, et al.
Published: (2024)
by: Su, Yukun, et al.
Published: (2024)
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
by: Zhang, Yunzhu, et al.
Published: (2025)
by: Zhang, Yunzhu, et al.
Published: (2025)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
by: He, Haibin, et al.
Published: (2026)
by: He, Haibin, et al.
Published: (2026)
Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings
by: Wu, Peixi, et al.
Published: (2026)
by: Wu, Peixi, et al.
Published: (2026)
From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment
by: Suo, Yucheng, et al.
Published: (2025)
by: Suo, Yucheng, et al.
Published: (2025)
HandRefiner: Refining Malformed Hands in Generated Images by Diffusion-based Conditional Inpainting
by: Lu, Wenquan, et al.
Published: (2023)
by: Lu, Wenquan, et al.
Published: (2023)
WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
by: Tang, Changli, et al.
Published: (2025)
by: Tang, Changli, et al.
Published: (2025)
On the Element-Wise Representation and Reasoning in Zero-Shot Image Recognition: A Systematic Survey
by: Guo, Jingcai, et al.
Published: (2024)
by: Guo, Jingcai, et al.
Published: (2024)
CompAgent: An Agentic Framework for Visual Compliance Verification
by: Ghosh, Rahul, et al.
Published: (2025)
by: Ghosh, Rahul, et al.
Published: (2025)
TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning
by: Pan, Junwen, et al.
Published: (2025)
by: Pan, Junwen, et al.
Published: (2025)
V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval
by: Chen, Dongyang, et al.
Published: (2026)
by: Chen, Dongyang, et al.
Published: (2026)
Towards Autonomous UAV Visual Object Search in City Space: Benchmark and Agentic Methodology
by: Ji, Yatai, et al.
Published: (2025)
by: Ji, Yatai, et al.
Published: (2025)
Echo-α: Large Agentic Multimodal Reasoning Model for Ultrasound Interpretation
by: Zhang, Jing, et al.
Published: (2026)
by: Zhang, Jing, et al.
Published: (2026)
Cycle Context Verification for In-Context Medical Image Segmentation
by: Hu, Shishuai, et al.
Published: (2025)
by: Hu, Shishuai, et al.
Published: (2025)
Towards Long-horizon Agentic Multimodal Search
by: Du, Yifan, et al.
Published: (2026)
by: Du, Yifan, et al.
Published: (2026)
LongVidSearch: An Agentic Benchmark for Multi-hop Evidence Retrieval Planning in Long Videos
by: Yu, Rongyi, et al.
Published: (2026)
by: Yu, Rongyi, et al.
Published: (2026)
Similar Items
-
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
by: Yang, Jie, et al.
Published: (2025) -
SAIL: Self-Amplified Iterative Learning for Diffusion Model Alignment with Minimal Human Feedback
by: He, Xiaoxuan, et al.
Published: (2026) -
WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
by: Yang, Jian, et al.
Published: (2025) -
D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning
by: Tang, Changli, et al.
Published: (2026) -
Gen-Searcher: Reinforcing Agentic Search for Image Generation
by: Feng, Kaituo, et al.
Published: (2026)