LaVIDE: A Language-Vision Discriminator for Detecting Changes in Satellite Image with Map References
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Shuguo, Xu, Fang, Jia, Sen, Xia, Gui-Song |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DMTG: One-Shot Differentiable Multi-Task Grouping
by: Gao, Yuan, et al.
Published: (2024)
by: Gao, Yuan, et al.
Published: (2024)
MetaDent: Labeling Clinical Images for Vision-Language Models in Dentistry
by: Li, Meng-Xun, et al.
Published: (2026)
by: Li, Meng-Xun, et al.
Published: (2026)
Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image
by: Qian, Ming, et al.
Published: (2026)
by: Qian, Ming, et al.
Published: (2026)
Refer to Any Segmentation Mask Group With Vision-Language Prompts
by: Cao, Shengcao, et al.
Published: (2025)
by: Cao, Shengcao, et al.
Published: (2025)
Fake-HR1: Rethinking Reasoning of Vision Language Model for Synthetic Image Detection
by: Jiang, Changjiang, et al.
Published: (2026)
by: Jiang, Changjiang, et al.
Published: (2026)
Class-Discriminative Attention Maps for Vision Transformers
by: Brocki, Lennart, et al.
Published: (2023)
by: Brocki, Lennart, et al.
Published: (2023)
MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection
by: Shi, Kuo, et al.
Published: (2025)
by: Shi, Kuo, et al.
Published: (2025)
AdPO: Enhancing the Adversarial Robustness of Large Vision-Language Models with Preference Optimization
by: Liu, Chaohu, et al.
Published: (2025)
by: Liu, Chaohu, et al.
Published: (2025)
HeatPrompt: Zero-Shot Vision-Language Modeling of Urban Heat Demand from Satellite Images
by: Thota, Kundan, et al.
Published: (2026)
by: Thota, Kundan, et al.
Published: (2026)
Towards Temporal Change Explanations from Bi-Temporal Satellite Images
by: Tsujimoto, Ryo, et al.
Published: (2024)
by: Tsujimoto, Ryo, et al.
Published: (2024)
A Structured Review of Underwater Object Detection Challenges and Solutions: From Traditional to Large Vision Language Models
by: Nabahirwa, Edwine, et al.
Published: (2025)
by: Nabahirwa, Edwine, et al.
Published: (2025)
Beyond Matching to Tiles: Bridging Unaligned Aerial and Satellite Views for Vision-Only UAV Navigation
by: Liu, Kejia, et al.
Published: (2026)
by: Liu, Kejia, et al.
Published: (2026)
Cross-aware Early Fusion with Stage-divided Vision and Language Transformer Encoders for Referring Image Segmentation
by: Cho, Yubin, et al.
Published: (2024)
by: Cho, Yubin, et al.
Published: (2024)
MapChange: Enhancing Semantic Change Detection with Temporal-Invariant Historical Maps Based on Deep Triplet Network
by: Liu, Yinhe, et al.
Published: (2024)
by: Liu, Yinhe, et al.
Published: (2024)
LL-ICM: Image Compression for Low-level Machine Vision via Large Vision-Language Model
by: Xue, Yuan, et al.
Published: (2024)
by: Xue, Yuan, et al.
Published: (2024)
RAU: Reference-based Anatomical Understanding with Vision Language Models
by: Li, Yiwei, et al.
Published: (2025)
by: Li, Yiwei, et al.
Published: (2025)
MonitorVLM:A Vision Language Framework for Safety Violation Detection in Mining Operations
by: Wu, Jiang, et al.
Published: (2025)
by: Wu, Jiang, et al.
Published: (2025)
On the Utility of Foundation Models for Fast MRI: Vision-Language-Guided Image Reconstruction
by: Feng, Ruimin, et al.
Published: (2025)
by: Feng, Ruimin, et al.
Published: (2025)
Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
by: Ye, Wei, et al.
Published: (2024)
by: Ye, Wei, et al.
Published: (2024)
DAE-Fuse: An Adaptive Discriminative Autoencoder for Multi-Modality Image Fusion
by: Guo, Yuchen, et al.
Published: (2024)
by: Guo, Yuchen, et al.
Published: (2024)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
TAMMs: Change Understanding and Forecasting in Satellite Image Time Series with Temporal-Aware Multimodal Models
by: Guo, Zhongbin, et al.
Published: (2025)
by: Guo, Zhongbin, et al.
Published: (2025)
CMEdataset Advancing China Map Detection and Standardization with Digital Image Resources
by: Xu, Yan, et al.
Published: (2025)
by: Xu, Yan, et al.
Published: (2025)
TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation
by: Chen, Jiaqi, et al.
Published: (2024)
by: Chen, Jiaqi, et al.
Published: (2024)
Rethinking Token Reduction for Large Vision-Language Models
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
Integrating Traditional and Deep Learning Methods to Detect Tree Crowns in Satellite Images
by: Durgut, Ozan, et al.
Published: (2025)
by: Durgut, Ozan, et al.
Published: (2025)
ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models
by: Yi, Jingwei, et al.
Published: (2025)
by: Yi, Jingwei, et al.
Published: (2025)
Bias Detection and Rotation-Robustness Mitigation in Vision-Language Models and Generative Image Models
by: Mithila, Tarannum
Published: (2026)
by: Mithila, Tarannum
Published: (2026)
ViLaCD-R1: A Vision-Language Framework for Semantic Change Detection in Remote Sensing
by: Ma, Xingwei, et al.
Published: (2025)
by: Ma, Xingwei, et al.
Published: (2025)
SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning
by: Wu, Xue, et al.
Published: (2026)
by: Wu, Xue, et al.
Published: (2026)
Complementing Onboard Sensors with Satellite Map: A New Perspective for HD Map Construction
by: Gao, Wenjie, et al.
Published: (2023)
by: Gao, Wenjie, et al.
Published: (2023)
Adversarial Prompt Injection Attack on Multimodal Large Language Models
by: Ding, Meiwen, et al.
Published: (2026)
by: Ding, Meiwen, et al.
Published: (2026)
NavOne: One-Step Global Planning for Vision-Language Navigation on Top-Down Maps
by: Zhan, Dijia, et al.
Published: (2026)
by: Zhan, Dijia, et al.
Published: (2026)
KKA: Improving Vision Anomaly Detection through Anomaly-related Knowledge from Large Language Models
by: Chen, Dong, et al.
Published: (2025)
by: Chen, Dong, et al.
Published: (2025)
Automatic Image Annotation for Mapped Features Detection
by: Noizet, Maxime, et al.
Published: (2024)
by: Noizet, Maxime, et al.
Published: (2024)
Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models
by: Liang, Shuang, et al.
Published: (2025)
by: Liang, Shuang, et al.
Published: (2025)
MapDream: Task-Driven Map Learning for Vision-Language Navigation
by: Lian, Guoxin, et al.
Published: (2026)
by: Lian, Guoxin, et al.
Published: (2026)
OmniOVCD: Streamlining Open-Vocabulary Change Detection with SAM 3
by: Zhang, Xu, et al.
Published: (2026)
by: Zhang, Xu, et al.
Published: (2026)
Similar Items
-
DMTG: One-Shot Differentiable Multi-Task Grouping
by: Gao, Yuan, et al.
Published: (2024) -
MetaDent: Labeling Clinical Images for Vision-Language Models in Dentistry
by: Li, Meng-Xun, et al.
Published: (2026) -
Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image
by: Qian, Ming, et al.
Published: (2026) -
Refer to Any Segmentation Mask Group With Vision-Language Prompts
by: Cao, Shengcao, et al.
Published: (2025) -
Fake-HR1: Rethinking Reasoning of Vision Language Model for Synthetic Image Detection
by: Jiang, Changjiang, et al.
Published: (2026)