SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Qiwei, Lu, Xukun, Liu, Wang, Duan, Puhong, Kang, Xudong, Li, Shutao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agriculture-Vision Challenge 2024 -- The Runner-Up Solution for Agricultural Pattern Recognition via Class Balancing and Model Ensemble
by: Liu, Wang, et al.
Published: (2024)
by: Liu, Wang, et al.
Published: (2024)
Learning from Noisy Pseudo-labels for All-Weather Land Cover Mapping
by: Liu, Wang, et al.
Published: (2025)
by: Liu, Wang, et al.
Published: (2025)
SAR-AE-SFP: SAR Imagery Adversarial Example in Real Physics domain with Target Scattering Feature Parameters
by: Cui, Jiahao, et al.
Published: (2024)
by: Cui, Jiahao, et al.
Published: (2024)
Prompting Foundation Models for Zero-Shot Ship Instance Segmentation in SAR Imagery
by: Mansour, Islam, et al.
Published: (2026)
by: Mansour, Islam, et al.
Published: (2026)
SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning
by: Wu, Xue, et al.
Published: (2026)
by: Wu, Xue, et al.
Published: (2026)
Landsat30-AU: A Vision-Language Dataset for Australian Landsat Imagery
by: Ma, Sai, et al.
Published: (2025)
by: Ma, Sai, et al.
Published: (2025)
FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery
by: Zhang, Xiaokun, et al.
Published: (2026)
by: Zhang, Xiaokun, et al.
Published: (2026)
Understanding the Transfer Limits of Vision Foundation Models
by: Huang, Shiqi, et al.
Published: (2026)
by: Huang, Shiqi, et al.
Published: (2026)
When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning
by: Luo, Junwei, et al.
Published: (2025)
by: Luo, Junwei, et al.
Published: (2025)
Examining the Commitments and Difficulties Inherent in Multimodal Foundation Models for Street View Imagery
by: Yang, Zhenyuan, et al.
Published: (2024)
by: Yang, Zhenyuan, et al.
Published: (2024)
When are Foundation Models Effective? Understanding the Suitability for Pixel-Level Classification Using Multispectral Imagery
by: Xie, Yiqun, et al.
Published: (2024)
by: Xie, Yiqun, et al.
Published: (2024)
SoccerMaster: A Vision Foundation Model for Soccer Understanding
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models
by: Ma, Jie, et al.
Published: (2026)
by: Ma, Jie, et al.
Published: (2026)
Smart Transfer: Leveraging Vision Foundation Model for Rapid Building Damage Mapping with Post-Earthquake VHR Imagery
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
Bootstrapping Diffusion: Diffusion Model Training Leveraging Partial and Corrupted Data
by: Ma, Xudong
Published: (2025)
by: Ma, Xudong
Published: (2025)
RAU: Reference-based Anatomical Understanding with Vision Language Models
by: Li, Yiwei, et al.
Published: (2025)
by: Li, Yiwei, et al.
Published: (2025)
Lightweight Multimodal Adaptation of Vision Language Models for Species Recognition and Habitat Context Interpretation in Drone Thermal Imagery
by: Chen, Hao, et al.
Published: (2026)
by: Chen, Hao, et al.
Published: (2026)
Effective Damage Data Generation by Fusing Imagery with Human Knowledge Using Vision-Language Models
by: Wei, Jie, et al.
Published: (2025)
by: Wei, Jie, et al.
Published: (2025)
Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models
by: Hu, Nanxing, et al.
Published: (2025)
by: Hu, Nanxing, et al.
Published: (2025)
All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker Adaptation
by: Wang, Xudong, et al.
Published: (2026)
by: Wang, Xudong, et al.
Published: (2026)
Deep Hashing with Semantic Hash Centers for Image Retrieval
by: Chen, Li, et al.
Published: (2025)
by: Chen, Li, et al.
Published: (2025)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
by: Liu, Hanqing, et al.
Published: (2026)
by: Liu, Hanqing, et al.
Published: (2026)
Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications
by: Rahman, Ben
Published: (2025)
by: Rahman, Ben
Published: (2025)
Heuristical Comparison of Vision Transformers Against Convolutional Neural Networks for Semantic Segmentation on Remote Sensing Imagery
by: Dahal, Ashim, et al.
Published: (2024)
by: Dahal, Ashim, et al.
Published: (2024)
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
by: Zhang, Jihai, et al.
Published: (2025)
by: Zhang, Jihai, et al.
Published: (2025)
Causality Model for Semantic Understanding on Videos
by: Yicong, Li
Published: (2025)
by: Yicong, Li
Published: (2025)
Multimodal Prompt Alignment for Facial Expression Recognition
by: Ma, Fuyan, et al.
Published: (2025)
by: Ma, Fuyan, et al.
Published: (2025)
Understanding Graphical Perception in Data Visualization through Zero-shot Prompting of Vision-Language Models
by: Guo, Grace, et al.
Published: (2024)
by: Guo, Grace, et al.
Published: (2024)
Reasoning under Vision: Understanding Visual-Spatial Cognition in Vision-Language Models for CAPTCHA
by: Song, Python, et al.
Published: (2025)
by: Song, Python, et al.
Published: (2025)
On the Utility of Foundation Models for Fast MRI: Vision-Language-Guided Image Reconstruction
by: Feng, Ruimin, et al.
Published: (2025)
by: Feng, Ruimin, et al.
Published: (2025)
FloodVision: Urban Flood Depth Estimation Using Foundation Vision-Language Models and Domain Knowledge Graph
by: Liu, Zhangding, et al.
Published: (2025)
by: Liu, Zhangding, et al.
Published: (2025)
Understanding Counting Mechanisms in Large Language and Vision-Language Models
by: Hasani, Hosein, et al.
Published: (2025)
by: Hasani, Hosein, et al.
Published: (2025)
From Isolated Islands to Pangea: Unifying Semantic Space for Human Action Understanding
by: Li, Yong-Lu, et al.
Published: (2023)
by: Li, Yong-Lu, et al.
Published: (2023)
A Visual Semantic Adaptive Watermark grounded by Prefix-Tuning for Large Vision-Language Model
by: Zheng, Qi, et al.
Published: (2026)
by: Zheng, Qi, et al.
Published: (2026)
Differential Privacy Image Generation with Reconstruction Loss and Noise Injection Using an Error Feedback SGD
by: Ma, Qiwei, et al.
Published: (2026)
by: Ma, Qiwei, et al.
Published: (2026)
Remodeling Semantic Relationships in Vision-Language Fine-Tuning
by: Wu, Xiangyang, et al.
Published: (2025)
by: Wu, Xiangyang, et al.
Published: (2025)
Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation
by: Shi, Jin, et al.
Published: (2026)
by: Shi, Jin, et al.
Published: (2026)
Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models
by: Li, Nanxi, et al.
Published: (2026)
by: Li, Nanxi, et al.
Published: (2026)
SatVision-TOA: A Geospatial Foundation Model for Coarse-Resolution All-Sky Remote Sensing Imagery
by: Spradlin, Caleb S., et al.
Published: (2024)
by: Spradlin, Caleb S., et al.
Published: (2024)
Merlin: A Computed Tomography Vision-Language Foundation Model and Dataset
by: Blankemeier, Louis, et al.
Published: (2024)
by: Blankemeier, Louis, et al.
Published: (2024)
Similar Items
-
Agriculture-Vision Challenge 2024 -- The Runner-Up Solution for Agricultural Pattern Recognition via Class Balancing and Model Ensemble
by: Liu, Wang, et al.
Published: (2024) -
Learning from Noisy Pseudo-labels for All-Weather Land Cover Mapping
by: Liu, Wang, et al.
Published: (2025) -
SAR-AE-SFP: SAR Imagery Adversarial Example in Real Physics domain with Target Scattering Feature Parameters
by: Cui, Jiahao, et al.
Published: (2024) -
Prompting Foundation Models for Zero-Shot Ship Instance Segmentation in SAR Imagery
by: Mansour, Islam, et al.
Published: (2026) -
SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning
by: Wu, Xue, et al.
Published: (2026)