Saved in:
| Main Authors: | Fuad, Nafis, Qian, Xiaodong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.17108 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FloodVision: Urban Flood Depth Estimation Using Foundation Vision-Language Models and Domain Knowledge Graph
by: Liu, Zhangding, et al.
Published: (2025)
by: Liu, Zhangding, et al.
Published: (2025)
Interpretable Vision Transformers in Monocular Depth Estimation via SVDA
by: Arampatzakis, Vasileios, et al.
Published: (2026)
by: Arampatzakis, Vasileios, et al.
Published: (2026)
From Local to Global to Mechanistic: An iERF-Centered Unified Framework for Interpreting Vision Models
by: Kim, Yearim, et al.
Published: (2026)
by: Kim, Yearim, et al.
Published: (2026)
Scale Alone Does not Improve Mechanistic Interpretability in Vision Models
by: Zimmermann, Roland S., et al.
Published: (2023)
by: Zimmermann, Roland S., et al.
Published: (2023)
Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
by: Li, Qiming, et al.
Published: (2025)
by: Li, Qiming, et al.
Published: (2025)
Vision-Language Embodiment for Monocular Depth Estimation
by: Zhang, Jinchang, et al.
Published: (2025)
by: Zhang, Jinchang, et al.
Published: (2025)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
by: Che, Liwei, et al.
Published: (2026)
by: Che, Liwei, et al.
Published: (2026)
Interpretable Debiasing of Vision-Language Models for Social Fairness
by: An, Na Min, et al.
Published: (2026)
by: An, Na Min, et al.
Published: (2026)
DepthLM: Metric Depth From Vision Language Models
by: Cai, Zhipeng, et al.
Published: (2025)
by: Cai, Zhipeng, et al.
Published: (2025)
Towards Depth Foundation Model: Recent Trends in Vision-Based Depth Estimation
by: Xu, Zhen, et al.
Published: (2025)
by: Xu, Zhen, et al.
Published: (2025)
Geometric Flood Depth Estimation: Fusing Transformer-Based Segmentation with Digital Elevation Models
by: Le, Nhut, et al.
Published: (2026)
by: Le, Nhut, et al.
Published: (2026)
Improving Interpretability of Deep Active Learning for Flood Inundation Mapping Through Class Ambiguity Indices Using Multi-spectral Satellite Imagery
by: Lee, Hyunho, et al.
Published: (2024)
by: Lee, Hyunho, et al.
Published: (2024)
Automated Floodwater Depth Estimation Using Large Multimodal Model for Rapid Flood Mapping
by: Akinboyewa, Temitope, et al.
Published: (2024)
by: Akinboyewa, Temitope, et al.
Published: (2024)
Enhancing Geo-localization for Crowdsourced Flood Imagery via LLM-Guided Attention
by: Xu, Fengyi, et al.
Published: (2025)
by: Xu, Fengyi, et al.
Published: (2025)
Lightweight Multimodal Adaptation of Vision Language Models for Species Recognition and Habitat Context Interpretation in Drone Thermal Imagery
by: Chen, Hao, et al.
Published: (2026)
by: Chen, Hao, et al.
Published: (2026)
Recov-Vision: Linking Street View Imagery and Vision-Language Models for Post-Disaster Recovery
by: Xiao, Yiming, et al.
Published: (2025)
by: Xiao, Yiming, et al.
Published: (2025)
Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
by: Lee, Phillip Y., et al.
Published: (2025)
by: Lee, Phillip Y., et al.
Published: (2025)
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
by: Stan, Gabriela Ben Melech, et al.
Published: (2024)
by: Stan, Gabriela Ben Melech, et al.
Published: (2024)
MMRL++: Parameter-Efficient and Interaction-Aware Representation Learning for Vision-Language Models
by: Guo, Yuncheng, et al.
Published: (2025)
by: Guo, Yuncheng, et al.
Published: (2025)
Language as Prior, Vision as Calibration: Metric Scale Recovery for Monocular Depth Estimation
by: Zhan, Mingxia, et al.
Published: (2026)
by: Zhan, Mingxia, et al.
Published: (2026)
RVLM: Recursive Vision-Language Models with Adaptive Depth
by: Mayumu, Nicanor, et al.
Published: (2026)
by: Mayumu, Nicanor, et al.
Published: (2026)
Monocular Depth Estimation with Global-Aware Discretization and Local Context Modeling
by: Wu, Heng, et al.
Published: (2025)
by: Wu, Heng, et al.
Published: (2025)
Assessing Building Heat Resilience Using UAV and Street-View Imagery with Coupled Global Context Vision Transformer
by: Knoblauch, Steffen, et al.
Published: (2026)
by: Knoblauch, Steffen, et al.
Published: (2026)
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
by: Yuan, Tianyuan, et al.
Published: (2025)
by: Yuan, Tianyuan, et al.
Published: (2025)
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
by: Chatterjee, Agneet, et al.
Published: (2024)
by: Chatterjee, Agneet, et al.
Published: (2024)
AIFloodSense: A Global Aerial Imagery Dataset for Semantic Segmentation and Understanding of Flooded Environments
by: Simantiris, Georgios, et al.
Published: (2025)
by: Simantiris, Georgios, et al.
Published: (2025)
KptLLM: Unveiling the Power of Large Language Model for Keypoint Comprehension
by: Yang, Jie, et al.
Published: (2024)
by: Yang, Jie, et al.
Published: (2024)
Remote Sensing Imagery for Flood Detection: Exploration of Augmentation Strategies
by: Polushko, Vladyslav, et al.
Published: (2025)
by: Polushko, Vladyslav, et al.
Published: (2025)
INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
by: Bagaria, Anshul
Published: (2025)
by: Bagaria, Anshul
Published: (2025)
Detecting Unauthorized Vehicles using Deep Learning for Smart Cities: A Case Study on Bangladesh
by: Sukanto, Sudipto Das, et al.
Published: (2025)
by: Sukanto, Sudipto Das, et al.
Published: (2025)
Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
by: Lin, Tao, et al.
Published: (2026)
by: Lin, Tao, et al.
Published: (2026)
Generalizable Prompt Tuning for Vision-Language Models
by: Zhang, Qian
Published: (2024)
by: Zhang, Qian
Published: (2024)
Depth Supervised Neural Surface Reconstruction from Airborne Imagery
by: Hackstein, Vincent, et al.
Published: (2024)
by: Hackstein, Vincent, et al.
Published: (2024)
Mechanistically Guided LoRA Improves Paraphrase Consistency in Medical Vision-Language Models
by: Sadanandan, Binesh, et al.
Published: (2026)
by: Sadanandan, Binesh, et al.
Published: (2026)
MMRL: Multi-Modal Representation Learning for Vision-Language Models
by: Guo, Yuncheng, et al.
Published: (2025)
by: Guo, Yuncheng, et al.
Published: (2025)
AffordanceLLM: Grounding Affordance from Vision Language Models
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
by: Wu, Shiwei, et al.
Published: (2024)
by: Wu, Shiwei, et al.
Published: (2024)
PIFF: A Physics-Informed Generative Flow Model for Real-Time Flood Depth Mapping
by: Wu, ChunLiang, et al.
Published: (2025)
by: Wu, ChunLiang, et al.
Published: (2025)
Few-Shot Vision-Language Reasoning for Satellite Imagery via Verifiable Rewards
by: Koksal, Aybora, et al.
Published: (2025)
by: Koksal, Aybora, et al.
Published: (2025)
Mechanistic Interpretability of Diffusion Models: Circuit-Level Analysis and Causal Validation
by: Roy, Dip
Published: (2025)
by: Roy, Dip
Published: (2025)
Similar Items
-
FloodVision: Urban Flood Depth Estimation Using Foundation Vision-Language Models and Domain Knowledge Graph
by: Liu, Zhangding, et al.
Published: (2025) -
Interpretable Vision Transformers in Monocular Depth Estimation via SVDA
by: Arampatzakis, Vasileios, et al.
Published: (2026) -
From Local to Global to Mechanistic: An iERF-Centered Unified Framework for Interpreting Vision Models
by: Kim, Yearim, et al.
Published: (2026) -
Scale Alone Does not Improve Mechanistic Interpretability in Vision Models
by: Zimmermann, Roland S., et al.
Published: (2023) -
Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
by: Li, Qiming, et al.
Published: (2025)