SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
Fuente:
arXiv
Salvato in:
| Autori principali: | Ma, Qiwei, Lu, Xukun, Liu, Wang, Duan, Puhong, Kang, Xudong, Li, Shutao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Agriculture-Vision Challenge 2024 -- The Runner-Up Solution for Agricultural Pattern Recognition via Class Balancing and Model Ensemble
di: Liu, Wang, et al.
Pubblicazione: (2024)
di: Liu, Wang, et al.
Pubblicazione: (2024)
Learning from Noisy Pseudo-labels for All-Weather Land Cover Mapping
di: Liu, Wang, et al.
Pubblicazione: (2025)
di: Liu, Wang, et al.
Pubblicazione: (2025)
SAR-AE-SFP: SAR Imagery Adversarial Example in Real Physics domain with Target Scattering Feature Parameters
di: Cui, Jiahao, et al.
Pubblicazione: (2024)
di: Cui, Jiahao, et al.
Pubblicazione: (2024)
Prompting Foundation Models for Zero-Shot Ship Instance Segmentation in SAR Imagery
di: Mansour, Islam, et al.
Pubblicazione: (2026)
di: Mansour, Islam, et al.
Pubblicazione: (2026)
SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning
di: Wu, Xue, et al.
Pubblicazione: (2026)
di: Wu, Xue, et al.
Pubblicazione: (2026)
Landsat30-AU: A Vision-Language Dataset for Australian Landsat Imagery
di: Ma, Sai, et al.
Pubblicazione: (2025)
di: Ma, Sai, et al.
Pubblicazione: (2025)
FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery
di: Zhang, Xiaokun, et al.
Pubblicazione: (2026)
di: Zhang, Xiaokun, et al.
Pubblicazione: (2026)
Understanding the Transfer Limits of Vision Foundation Models
di: Huang, Shiqi, et al.
Pubblicazione: (2026)
di: Huang, Shiqi, et al.
Pubblicazione: (2026)
When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning
di: Luo, Junwei, et al.
Pubblicazione: (2025)
di: Luo, Junwei, et al.
Pubblicazione: (2025)
Examining the Commitments and Difficulties Inherent in Multimodal Foundation Models for Street View Imagery
di: Yang, Zhenyuan, et al.
Pubblicazione: (2024)
di: Yang, Zhenyuan, et al.
Pubblicazione: (2024)
When are Foundation Models Effective? Understanding the Suitability for Pixel-Level Classification Using Multispectral Imagery
di: Xie, Yiqun, et al.
Pubblicazione: (2024)
di: Xie, Yiqun, et al.
Pubblicazione: (2024)
SoccerMaster: A Vision Foundation Model for Soccer Understanding
di: Yang, Haolin, et al.
Pubblicazione: (2025)
di: Yang, Haolin, et al.
Pubblicazione: (2025)
Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models
di: Ma, Jie, et al.
Pubblicazione: (2026)
di: Ma, Jie, et al.
Pubblicazione: (2026)
Smart Transfer: Leveraging Vision Foundation Model for Rapid Building Damage Mapping with Post-Earthquake VHR Imagery
di: Li, Hao, et al.
Pubblicazione: (2026)
di: Li, Hao, et al.
Pubblicazione: (2026)
Bootstrapping Diffusion: Diffusion Model Training Leveraging Partial and Corrupted Data
di: Ma, Xudong
Pubblicazione: (2025)
di: Ma, Xudong
Pubblicazione: (2025)
RAU: Reference-based Anatomical Understanding with Vision Language Models
di: Li, Yiwei, et al.
Pubblicazione: (2025)
di: Li, Yiwei, et al.
Pubblicazione: (2025)
Lightweight Multimodal Adaptation of Vision Language Models for Species Recognition and Habitat Context Interpretation in Drone Thermal Imagery
di: Chen, Hao, et al.
Pubblicazione: (2026)
di: Chen, Hao, et al.
Pubblicazione: (2026)
Effective Damage Data Generation by Fusing Imagery with Human Knowledge Using Vision-Language Models
di: Wei, Jie, et al.
Pubblicazione: (2025)
di: Wei, Jie, et al.
Pubblicazione: (2025)
Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models
di: Hu, Nanxing, et al.
Pubblicazione: (2025)
di: Hu, Nanxing, et al.
Pubblicazione: (2025)
All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker Adaptation
di: Wang, Xudong, et al.
Pubblicazione: (2026)
di: Wang, Xudong, et al.
Pubblicazione: (2026)
Deep Hashing with Semantic Hash Centers for Image Retrieval
di: Chen, Li, et al.
Pubblicazione: (2025)
di: Chen, Li, et al.
Pubblicazione: (2025)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
di: Liu, Hanqing, et al.
Pubblicazione: (2026)
di: Liu, Hanqing, et al.
Pubblicazione: (2026)
Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications
di: Rahman, Ben
Pubblicazione: (2025)
di: Rahman, Ben
Pubblicazione: (2025)
Heuristical Comparison of Vision Transformers Against Convolutional Neural Networks for Semantic Segmentation on Remote Sensing Imagery
di: Dahal, Ashim, et al.
Pubblicazione: (2024)
di: Dahal, Ashim, et al.
Pubblicazione: (2024)
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
di: Zhang, Jihai, et al.
Pubblicazione: (2025)
di: Zhang, Jihai, et al.
Pubblicazione: (2025)
Causality Model for Semantic Understanding on Videos
di: Yicong, Li
Pubblicazione: (2025)
di: Yicong, Li
Pubblicazione: (2025)
Multimodal Prompt Alignment for Facial Expression Recognition
di: Ma, Fuyan, et al.
Pubblicazione: (2025)
di: Ma, Fuyan, et al.
Pubblicazione: (2025)
Understanding Graphical Perception in Data Visualization through Zero-shot Prompting of Vision-Language Models
di: Guo, Grace, et al.
Pubblicazione: (2024)
di: Guo, Grace, et al.
Pubblicazione: (2024)
Reasoning under Vision: Understanding Visual-Spatial Cognition in Vision-Language Models for CAPTCHA
di: Song, Python, et al.
Pubblicazione: (2025)
di: Song, Python, et al.
Pubblicazione: (2025)
On the Utility of Foundation Models for Fast MRI: Vision-Language-Guided Image Reconstruction
di: Feng, Ruimin, et al.
Pubblicazione: (2025)
di: Feng, Ruimin, et al.
Pubblicazione: (2025)
FloodVision: Urban Flood Depth Estimation Using Foundation Vision-Language Models and Domain Knowledge Graph
di: Liu, Zhangding, et al.
Pubblicazione: (2025)
di: Liu, Zhangding, et al.
Pubblicazione: (2025)
Understanding Counting Mechanisms in Large Language and Vision-Language Models
di: Hasani, Hosein, et al.
Pubblicazione: (2025)
di: Hasani, Hosein, et al.
Pubblicazione: (2025)
From Isolated Islands to Pangea: Unifying Semantic Space for Human Action Understanding
di: Li, Yong-Lu, et al.
Pubblicazione: (2023)
di: Li, Yong-Lu, et al.
Pubblicazione: (2023)
A Visual Semantic Adaptive Watermark grounded by Prefix-Tuning for Large Vision-Language Model
di: Zheng, Qi, et al.
Pubblicazione: (2026)
di: Zheng, Qi, et al.
Pubblicazione: (2026)
Differential Privacy Image Generation with Reconstruction Loss and Noise Injection Using an Error Feedback SGD
di: Ma, Qiwei, et al.
Pubblicazione: (2026)
di: Ma, Qiwei, et al.
Pubblicazione: (2026)
Remodeling Semantic Relationships in Vision-Language Fine-Tuning
di: Wu, Xiangyang, et al.
Pubblicazione: (2025)
di: Wu, Xiangyang, et al.
Pubblicazione: (2025)
Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation
di: Shi, Jin, et al.
Pubblicazione: (2026)
di: Shi, Jin, et al.
Pubblicazione: (2026)
Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models
di: Li, Nanxi, et al.
Pubblicazione: (2026)
di: Li, Nanxi, et al.
Pubblicazione: (2026)
SatVision-TOA: A Geospatial Foundation Model for Coarse-Resolution All-Sky Remote Sensing Imagery
di: Spradlin, Caleb S., et al.
Pubblicazione: (2024)
di: Spradlin, Caleb S., et al.
Pubblicazione: (2024)
Merlin: A Computed Tomography Vision-Language Foundation Model and Dataset
di: Blankemeier, Louis, et al.
Pubblicazione: (2024)
di: Blankemeier, Louis, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Agriculture-Vision Challenge 2024 -- The Runner-Up Solution for Agricultural Pattern Recognition via Class Balancing and Model Ensemble
di: Liu, Wang, et al.
Pubblicazione: (2024) -
Learning from Noisy Pseudo-labels for All-Weather Land Cover Mapping
di: Liu, Wang, et al.
Pubblicazione: (2025) -
SAR-AE-SFP: SAR Imagery Adversarial Example in Real Physics domain with Target Scattering Feature Parameters
di: Cui, Jiahao, et al.
Pubblicazione: (2024) -
Prompting Foundation Models for Zero-Shot Ship Instance Segmentation in SAR Imagery
di: Mansour, Islam, et al.
Pubblicazione: (2026) -
SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning
di: Wu, Xue, et al.
Pubblicazione: (2026)