Where am I? Cross-View Geo-localization with Natural Language Descriptions
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Junyan, Lin, Honglin, Ou, Leyan, Chen, Dairong, Wang, Zihao, Zhu, Qi, He, Conghui, Li, Weijia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Street View Synthesis
by: Li, Weijia, et al.
Published: (2024)
by: Li, Weijia, et al.
Published: (2024)
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
by: Zhou, Baichuan, et al.
Published: (2024)
by: Zhou, Baichuan, et al.
Published: (2024)
SG-BEV: Satellite-Guided BEV Fusion for Cross-View Semantic Segmentation
by: Ye, Junyan, et al.
Published: (2024)
by: Ye, Junyan, et al.
Published: (2024)
Cross-view image geo-localization with Panorama-BEV Co-Retrieval Network
by: Ye, Junyan, et al.
Published: (2024)
by: Ye, Junyan, et al.
Published: (2024)
Fine-Grained Building Function Recognition from Street-View Images via Geometry-Aware Semi-Supervised Learning
by: Li, Weijia, et al.
Published: (2024)
by: Li, Weijia, et al.
Published: (2024)
FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection
by: Zhu, Leqi, et al.
Published: (2026)
by: Zhu, Leqi, et al.
Published: (2026)
"Where am I?" Scene Retrieval with Language
by: Chen, Jiaqi, et al.
Published: (2024)
by: Chen, Jiaqi, et al.
Published: (2024)
Leveraging BEV Paradigm for Ground-to-Aerial Image Synthesis
by: Ye, Junyan, et al.
Published: (2024)
by: Ye, Junyan, et al.
Published: (2024)
OmniAID: Decoupling Semantic and Artifacts for Universal AI-Generated Image Detection in the Wild
by: Guo, Yuncheng, et al.
Published: (2025)
by: Guo, Yuncheng, et al.
Published: (2025)
LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models
by: Ye, Junyan, et al.
Published: (2024)
by: Ye, Junyan, et al.
Published: (2024)
GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation
by: Yan, Zhiyuan, et al.
Published: (2025)
by: Yan, Zhiyuan, et al.
Published: (2025)
BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
by: Ye, Junyan, et al.
Published: (2025)
by: Ye, Junyan, et al.
Published: (2025)
Where Am I and What Will I See: An Auto-Regressive Model for Spatial Localization and View Prediction
by: Chen, Junyi, et al.
Published: (2024)
by: Chen, Junyi, et al.
Published: (2024)
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
by: Ye, Junyan, et al.
Published: (2025)
by: Ye, Junyan, et al.
Published: (2025)
RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
by: Ye, Junyan, et al.
Published: (2025)
by: Ye, Junyan, et al.
Published: (2025)
LEGION: Learning to Ground and Explain for Synthetic Image Detection
by: Kang, Hengrui, et al.
Published: (2025)
by: Kang, Hengrui, et al.
Published: (2025)
MOGeo: Beyond One-to-One Cross-View Object Geo-localization
by: Lv, Bo, et al.
Published: (2026)
by: Lv, Bo, et al.
Published: (2026)
Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact Explanation
by: Wen, Siwei, et al.
Published: (2025)
by: Wen, Siwei, et al.
Published: (2025)
Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?
by: Wen, Zichen, et al.
Published: (2025)
by: Wen, Zichen, et al.
Published: (2025)
Cross-View Consistency Regularisation for Knowledge Distillation
by: Zhang, Weijia, et al.
Published: (2024)
by: Zhang, Weijia, et al.
Published: (2024)
Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents
by: Feng, Peilin, et al.
Published: (2025)
by: Feng, Peilin, et al.
Published: (2025)
UrbanFeel: A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human Perspective
by: He, Jun, et al.
Published: (2025)
by: He, Jun, et al.
Published: (2025)
ConGeo: Robust Cross-view Geo-localization across Ground View Variations
by: Mi, Li, et al.
Published: (2024)
by: Mi, Li, et al.
Published: (2024)
MMFineReason: Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
by: Lin, Honglin, et al.
Published: (2026)
by: Lin, Honglin, et al.
Published: (2026)
GenClaw: Code-Driven Agentic Image Generation
by: Ye, Junyan, et al.
Published: (2026)
by: Ye, Junyan, et al.
Published: (2026)
Scene4U: Hierarchical Layered 3D Scene Reconstruction from Single Panoramic Image for Your Immerse Exploration
by: Huang, Zilong, et al.
Published: (2025)
by: Huang, Zilong, et al.
Published: (2025)
MajutsuCity: Language-driven Aesthetic-adaptive City Generation with Controllable 3D Assets and Layouts
by: Huang, Zilong, et al.
Published: (2025)
by: Huang, Zilong, et al.
Published: (2025)
GeoReasoner: Geo-localization with Reasoning in Street Views using a Large Vision-Language Model
by: Li, Ling, et al.
Published: (2024)
by: Li, Ling, et al.
Published: (2024)
UniABG: Unified Adversarial View Bridging and Graph Correspondence for Unsupervised Cross-View Geo-Localization
by: Chen, Cuiqun, et al.
Published: (2025)
by: Chen, Cuiqun, et al.
Published: (2025)
SatSAM2: Motion-Constrained Video Object Tracking in Satellite Imagery using Promptable SAM2 and Kalman Priors
by: Fan, Ruijie, et al.
Published: (2025)
by: Fan, Ruijie, et al.
Published: (2025)
Mind-Brush: Integrating Agentic Cognitive Search and Reasoning into Image Generation
by: He, Jun, et al.
Published: (2026)
by: He, Jun, et al.
Published: (2026)
DualGeo: A Dual-View Framework for Worldwide Image Geo-localization
by: Cui, Junchao, et al.
Published: (2026)
by: Cui, Junchao, et al.
Published: (2026)
NuNext: Reframing Nucleus Detection as Next-Point Detection
by: Shui, Zhongyi, et al.
Published: (2026)
by: Shui, Zhongyi, et al.
Published: (2026)
3D Building Reconstruction from Monocular Remote Sensing Images with Multi-level Supervisions
by: Li, Weijia, et al.
Published: (2024)
by: Li, Weijia, et al.
Published: (2024)
From Horizontal to Rotated: Cross-View Object Geo-Localization with Orientation Awareness
by: Fu, Chenlin, et al.
Published: (2026)
by: Fu, Chenlin, et al.
Published: (2026)
Parrot Captions Teach CLIP to Spot Text
by: Lin, Yiqi, et al.
Published: (2023)
by: Lin, Yiqi, et al.
Published: (2023)
Context-Aware Integration of Language and Visual References for Natural Language Tracking
by: Shao, Yanyan, et al.
Published: (2024)
by: Shao, Yanyan, et al.
Published: (2024)
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding
by: Liu, Zheng, et al.
Published: (2025)
by: Liu, Zheng, et al.
Published: (2025)
Window-to-Window BEV Representation Learning for Limited FoV Cross-View Geo-localization
by: Cheng, Lei, et al.
Published: (2024)
by: Cheng, Lei, et al.
Published: (2024)
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
by: Wen, Zichen, et al.
Published: (2025)
by: Wen, Zichen, et al.
Published: (2025)
Similar Items
-
CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Street View Synthesis
by: Li, Weijia, et al.
Published: (2024) -
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
by: Zhou, Baichuan, et al.
Published: (2024) -
SG-BEV: Satellite-Guided BEV Fusion for Cross-View Semantic Segmentation
by: Ye, Junyan, et al.
Published: (2024) -
Cross-view image geo-localization with Panorama-BEV Co-Retrieval Network
by: Ye, Junyan, et al.
Published: (2024) -
Fine-Grained Building Function Recognition from Street-View Images via Geometry-Aware Semi-Supervised Learning
by: Li, Weijia, et al.
Published: (2024)