AddressCLIP: Empowering Vision-Language Models for City-wide Image Address Localization
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Shixiong, Zhang, Chenghao, Fan, Lubin, Meng, Gaofeng, Xiang, Shiming, Ye, Jieping |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
by: Xu, Shixiong, et al.
Published: (2025)
by: Xu, Shixiong, et al.
Published: (2025)
Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models Stronger
by: Yang, Qi, et al.
Published: (2025)
by: Yang, Qi, et al.
Published: (2025)
Defying Imbalanced Forgetting in Class Incremental Learning
by: Xu, Shixiong, et al.
Published: (2024)
by: Xu, Shixiong, et al.
Published: (2024)
A Survey of Low-shot Vision-Language Model Adaptation via Representer Theorem
by: Ding, Kun, et al.
Published: (2024)
by: Ding, Kun, et al.
Published: (2024)
Calibrated Cache Model for Few-Shot Vision-Language Model Adaptation
by: Ding, Kun, et al.
Published: (2024)
by: Ding, Kun, et al.
Published: (2024)
CoL3D: Collaborative Learning of Single-view Depth and Camera Intrinsics for Metric 3D Shape Recovery
by: Zhang, Chenghao, et al.
Published: (2025)
by: Zhang, Chenghao, et al.
Published: (2025)
Enhancing Visual Continual Learning with Language-Guided Supervision
by: Ni, Bolin, et al.
Published: (2024)
by: Ni, Bolin, et al.
Published: (2024)
Reusable Architecture Growth for Continual Stereo Matching
by: Zhang, Chenghao, et al.
Published: (2024)
by: Zhang, Chenghao, et al.
Published: (2024)
Learning Neural Volumetric Pose Features for Camera Localization
by: Lin, Jingyu, et al.
Published: (2024)
by: Lin, Jingyu, et al.
Published: (2024)
SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
by: Chen, Pingyi, et al.
Published: (2025)
by: Chen, Pingyi, et al.
Published: (2025)
CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering
by: Hong, Yuyang, et al.
Published: (2026)
by: Hong, Yuyang, et al.
Published: (2026)
MSCI: Addressing CLIP's Inherent Limitations for Compositional Zero-Shot Learning
by: Wang, Yue, et al.
Published: (2025)
by: Wang, Yue, et al.
Published: (2025)
IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding
by: Zhu, Lanyun, et al.
Published: (2024)
by: Zhu, Lanyun, et al.
Published: (2024)
Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
by: Hong, Yuyang, et al.
Published: (2025)
by: Hong, Yuyang, et al.
Published: (2025)
EvoVLMA: Evolutionary Vision-Language Model Adaptation
by: Ding, Kun, et al.
Published: (2025)
by: Ding, Kun, et al.
Published: (2025)
Free Lunch for Generating Effective Outlier Supervision
by: Pei, Sen, et al.
Published: (2023)
by: Pei, Sen, et al.
Published: (2023)
NoPe-NeRF++: Local-to-Global Optimization of NeRF with No Pose Prior
by: Shi, Dongbo, et al.
Published: (2025)
by: Shi, Dongbo, et al.
Published: (2025)
PTZ-Calib: Robust Pan-Tilt-Zoom Camera Calibration
by: Guo, Jinhui, et al.
Published: (2025)
by: Guo, Jinhui, et al.
Published: (2025)
RemoteCLIP: A Vision Language Foundation Model for Remote Sensing
by: Liu, Fan, et al.
Published: (2023)
by: Liu, Fan, et al.
Published: (2023)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
by: Chen, Shiming, et al.
Published: (2025)
by: Chen, Shiming, et al.
Published: (2025)
Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature
by: Shen, Lingdong, et al.
Published: (2024)
by: Shen, Lingdong, et al.
Published: (2024)
Robust and Explainable Framework to Address Data Scarcity in Diagnostic Imaging
by: Zhao, Zehui, et al.
Published: (2024)
by: Zhao, Zehui, et al.
Published: (2024)
MolCLIP: A Molecular-Auxiliary CLIP Framework for Identifying Drug Mechanism of Action Based on Time-Lapsed Mitochondrial Images
by: Pang, Fengqian, et al.
Published: (2025)
by: Pang, Fengqian, et al.
Published: (2025)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
by: Huang, Yuchen, et al.
Published: (2025)
by: Huang, Yuchen, et al.
Published: (2025)
CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging
by: Imam, Raza, et al.
Published: (2024)
by: Imam, Raza, et al.
Published: (2024)
SocialFusion: Addressing Social Degradation in Pre-trained Vision-Language Models
by: Tahboub, Hamza, et al.
Published: (2025)
by: Tahboub, Hamza, et al.
Published: (2025)
Dual-Anchoring: Addressing State Drift in Vision-Language Navigation
by: Wu, Kangyi, et al.
Published: (2026)
by: Wu, Kangyi, et al.
Published: (2026)
Compositional Kronecker Context Optimization for Vision-Language Models
by: Ding, Kun, et al.
Published: (2024)
by: Ding, Kun, et al.
Published: (2024)
Morphological Addressing of Identity Basins in Text-to-Image Diffusion Models
by: Fraser, Andrew
Published: (2026)
by: Fraser, Andrew
Published: (2026)
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
by: Wei, Zhixiang, et al.
Published: (2025)
by: Wei, Zhixiang, et al.
Published: (2025)
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
by: Tong, Jintao, et al.
Published: (2025)
by: Tong, Jintao, et al.
Published: (2025)
ET tu, CLIP? Addressing Common Object Errors for Unseen Environments
by: Byun, Ye Won, et al.
Published: (2024)
by: Byun, Ye Won, et al.
Published: (2024)
Addressing Overthinking in Large Vision-Language Models via Gated Perception-Reasoning Optimization
by: Diao, Xingjian, et al.
Published: (2026)
by: Diao, Xingjian, et al.
Published: (2026)
DOFA-CLIP: Multimodal Vision-Language Foundation Models for Earth Observation
by: Xiong, Zhitong, et al.
Published: (2025)
by: Xiong, Zhitong, et al.
Published: (2025)
Diffuse-UDA: Addressing Unsupervised Domain Adaptation in Medical Image Segmentation with Appearance and Structure Aligned Diffusion Models
by: Gong, Haifan, et al.
Published: (2024)
by: Gong, Haifan, et al.
Published: (2024)
RWKV-CLIP: A Robust Vision-Language Representation Learner
by: Gu, Tiancheng, et al.
Published: (2024)
by: Gu, Tiancheng, et al.
Published: (2024)
ChatEarthNet: A Global-Scale Image-Text Dataset Empowering Vision-Language Geo-Foundation Models
by: Yuan, Zhenghang, et al.
Published: (2024)
by: Yuan, Zhenghang, et al.
Published: (2024)
Addressing Domain Discrepancy: A Dual-branch Collaborative Model to Unsupervised Dehazing
by: Fan, Shuaibin, et al.
Published: (2024)
by: Fan, Shuaibin, et al.
Published: (2024)
IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts
by: Wang, Juan, et al.
Published: (2026)
by: Wang, Juan, et al.
Published: (2026)
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
Similar Items
-
AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
by: Xu, Shixiong, et al.
Published: (2025) -
Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models Stronger
by: Yang, Qi, et al.
Published: (2025) -
Defying Imbalanced Forgetting in Class Incremental Learning
by: Xu, Shixiong, et al.
Published: (2024) -
A Survey of Low-shot Vision-Language Model Adaptation via Representer Theorem
by: Ding, Kun, et al.
Published: (2024) -
Calibrated Cache Model for Few-Shot Vision-Language Model Adaptation
by: Ding, Kun, et al.
Published: (2024)