Visual Geo-Localization from images
Fuente:
arXiv
Guardado en:
| Autores principales: | Saoud, Rania, Larabi, Slimane |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
More Consideration for the Perceptron
por: Larabi, Slimane
Publicado: (2024)
por: Larabi, Slimane
Publicado: (2024)
Uterine Ultrasound Image Captioning Using Deep Learning Techniques
por: Boulesnane, Abdennour, et al.
Publicado: (2024)
por: Boulesnane, Abdennour, et al.
Publicado: (2024)
Real-Time Threaded Houbara Detection and Segmentation for Wildlife Conservation using Mobile Platforms
por: Saoud, Lyes Saad, et al.
Publicado: (2025)
por: Saoud, Lyes Saad, et al.
Publicado: (2025)
Can Mental Imagery Improve the Thinking Capabilities of AI Systems?
por: Larabi, Slimane
Publicado: (2025)
por: Larabi, Slimane
Publicado: (2025)
From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning
por: Jiao, Pengkun, et al.
Publicado: (2024)
por: Jiao, Pengkun, et al.
Publicado: (2024)
RELO: Reinforcement Learning to Localize for Visual Object Tracking
por: Chen, Xin, et al.
Publicado: (2026)
por: Chen, Xin, et al.
Publicado: (2026)
DuoFormer: Leveraging Hierarchical Visual Representations by Local and Global Attention
por: Tang, Xiaoya, et al.
Publicado: (2024)
por: Tang, Xiaoya, et al.
Publicado: (2024)
LocalMamba: Visual State Space Model with Windowed Selective Scan
por: Huang, Tao, et al.
Publicado: (2024)
por: Huang, Tao, et al.
Publicado: (2024)
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
por: Jiang, Yubo, et al.
Publicado: (2026)
por: Jiang, Yubo, et al.
Publicado: (2026)
Visual Self-paced Iterative Learning for Unsupervised Temporal Action Localization
por: Hu, Yupeng, et al.
Publicado: (2023)
por: Hu, Yupeng, et al.
Publicado: (2023)
GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal Models
por: Zheng, Yushuo, et al.
Publicado: (2025)
por: Zheng, Yushuo, et al.
Publicado: (2025)
Just Zoom In: Cross-View Geo-Localization via Autoregressive Zooming
por: Erzurumlu, Yunus Talha, et al.
Publicado: (2026)
por: Erzurumlu, Yunus Talha, et al.
Publicado: (2026)
GLFNET: Global-Local (frequency) Filter Networks for efficient medical image segmentation
por: Tragakis, Athanasios, et al.
Publicado: (2024)
por: Tragakis, Athanasios, et al.
Publicado: (2024)
Locatability-Guided Adaptive Reasoning for Image Geo-Localization with Vision-Language Models
por: Yu, Bo, et al.
Publicado: (2026)
por: Yu, Bo, et al.
Publicado: (2026)
Dual-level Progressive Hardness-Aware Reweighting for Cross-View Geo-Localization
por: Zheng, Guozheng, et al.
Publicado: (2025)
por: Zheng, Guozheng, et al.
Publicado: (2025)
A Parameter-Efficient Mixture-of-Experts Framework for Cross-Modal Geo-Localization
por: Li, LinFeng, et al.
Publicado: (2025)
por: Li, LinFeng, et al.
Publicado: (2025)
Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization
por: Chen, Yuqi, et al.
Publicado: (2026)
por: Chen, Yuqi, et al.
Publicado: (2026)
SINA: A Circuit Schematic Image-to-Netlist Generator Using Artificial Intelligence
por: Aldowaish, Saoud, et al.
Publicado: (2026)
por: Aldowaish, Saoud, et al.
Publicado: (2026)
Semi-Supervised Audio-Visual Video Action Recognition with Audio Source Localization Guided Mixup
por: Kang, Seokun, et al.
Publicado: (2025)
por: Kang, Seokun, et al.
Publicado: (2025)
Open-Vocabulary Action Localization with Iterative Visual Prompting
por: Wake, Naoki, et al.
Publicado: (2024)
por: Wake, Naoki, et al.
Publicado: (2024)
Scaling Image Tokenizers with Grouped Spherical Quantization
por: Wang, Jiangtao, et al.
Publicado: (2024)
por: Wang, Jiangtao, et al.
Publicado: (2024)
Localization, balance and affinity: a stronger multifaceted collaborative salient object detector in remote sensing images
por: Xie, Yakun, et al.
Publicado: (2024)
por: Xie, Yakun, et al.
Publicado: (2024)
GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models
por: Zheng, Shurong, et al.
Publicado: (2026)
por: Zheng, Shurong, et al.
Publicado: (2026)
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
por: Fu, Ling, et al.
Publicado: (2024)
por: Fu, Ling, et al.
Publicado: (2024)
Swarm Intelligence in Geo-Localization: A Multi-Agent Large Vision-Language Model Collaborative Framework
por: Han, Xiao, et al.
Publicado: (2024)
por: Han, Xiao, et al.
Publicado: (2024)
Without Paired Labeled Data: End-to-End Self-Supervised Learning for Drone-view Geo-Localization
por: Chen, Zhongwei, et al.
Publicado: (2025)
por: Chen, Zhongwei, et al.
Publicado: (2025)
Rethinking 3D Dense Caption and Visual Grounding in A Unified Framework through Prompt-based Localization
por: Luo, Yongdong, et al.
Publicado: (2024)
por: Luo, Yongdong, et al.
Publicado: (2024)
Enhancing Adversarial Transferability in Visual-Language Pre-training Models via Local Shuffle and Sample-based Attack
por: Liu, Xin, et al.
Publicado: (2025)
por: Liu, Xin, et al.
Publicado: (2025)
GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imagery
por: Wang, Fengxiang, et al.
Publicado: (2026)
por: Wang, Fengxiang, et al.
Publicado: (2026)
AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
por: Zeng, Kang, et al.
Publicado: (2025)
por: Zeng, Kang, et al.
Publicado: (2025)
Beyond Static Cropping: Layer-Adaptive Visual Localization and Decoding Enhancement
por: Zhu, Zipeng, et al.
Publicado: (2026)
por: Zhu, Zipeng, et al.
Publicado: (2026)
Re-identification from histopathology images
por: Ganz, Jonathan, et al.
Publicado: (2024)
por: Ganz, Jonathan, et al.
Publicado: (2024)
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
por: Seyfioglu, Mehmet Saygin, et al.
Publicado: (2023)
por: Seyfioglu, Mehmet Saygin, et al.
Publicado: (2023)
Visual-ERM: Reward Modeling for Visual Equivalence
por: Liu, Ziyu, et al.
Publicado: (2026)
por: Liu, Ziyu, et al.
Publicado: (2026)
Visual Position Prompt for MLLM based Visual Grounding
por: Tang, Wei, et al.
Publicado: (2025)
por: Tang, Wei, et al.
Publicado: (2025)
Towards Understanding Visual Grounding in Visual Language Models
por: Pantazopoulos, Georgios, et al.
Publicado: (2025)
por: Pantazopoulos, Georgios, et al.
Publicado: (2025)
Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning For Visual Story Synthesis
por: Song, Tianyi, et al.
Publicado: (2023)
por: Song, Tianyi, et al.
Publicado: (2023)
Beyond Visual Memory: Mechanistic Diagnostics of Latent Visual Reasoning
por: Guo, Garvin, et al.
Publicado: (2026)
por: Guo, Garvin, et al.
Publicado: (2026)
Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension
por: Xu, Haoran, et al.
Publicado: (2026)
por: Xu, Haoran, et al.
Publicado: (2026)
Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery
por: Tran, Chi-Nguyen, et al.
Publicado: (2026)
por: Tran, Chi-Nguyen, et al.
Publicado: (2026)
Ejemplares similares
-
More Consideration for the Perceptron
por: Larabi, Slimane
Publicado: (2024) -
Uterine Ultrasound Image Captioning Using Deep Learning Techniques
por: Boulesnane, Abdennour, et al.
Publicado: (2024) -
Real-Time Threaded Houbara Detection and Segmentation for Wildlife Conservation using Mobile Platforms
por: Saoud, Lyes Saad, et al.
Publicado: (2025) -
Can Mental Imagery Improve the Thinking Capabilities of AI Systems?
por: Larabi, Slimane
Publicado: (2025) -
From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning
por: Jiao, Pengkun, et al.
Publicado: (2024)