GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Elgendy, Hosam, Sharshar, Ahmed, Aboeitta, Ahmed, Ashraf, Yasser, Guizani, Mohsen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ChatENV: An Interactive Vision-Language Model for Sensor-Guided Environmental Monitoring and Scenario Simulation
von: Elgendy, Hosam, et al.
Veröffentlicht: (2025)
von: Elgendy, Hosam, et al.
Veröffentlicht: (2025)
Not Only Grey Matter: OmniBrain for Robust Multimodal Classification of Alzheimer's Disease
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2025)
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2025)
PulmoFusion: Advancing Pulmonary Health with Efficient Multi-Modal Fusion
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2025)
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2025)
GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
von: Wang, Fengxiang, et al.
Veröffentlicht: (2025)
von: Wang, Fengxiang, et al.
Veröffentlicht: (2025)
Vision-Language Models for Edge Networks: A Comprehensive Survey
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2025)
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2025)
SPACT18: Spiking Human Action Recognition Benchmark Dataset with Complementary RGB and Thermal Modalities
von: Ashraf, Yasser, et al.
Veröffentlicht: (2025)
von: Ashraf, Yasser, et al.
Veröffentlicht: (2025)
Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2026)
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2026)
FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants
von: Bhosale, Mahesh, et al.
Veröffentlicht: (2026)
von: Bhosale, Mahesh, et al.
Veröffentlicht: (2026)
PeftCD: Leveraging Vision Foundation Models with Parameter-Efficient Fine-Tuning for Remote Sensing Change Detection
von: Dong, Sijun, et al.
Veröffentlicht: (2025)
von: Dong, Sijun, et al.
Veröffentlicht: (2025)
CityLLaVA: Efficient Fine-Tuning for VLMs in City Scenario
von: Duan, Zhizhao, et al.
Veröffentlicht: (2024)
von: Duan, Zhizhao, et al.
Veröffentlicht: (2024)
Constraint-Driven Warm-Freeze for Efficient Transfer Learning in Photovoltaic Systems
von: Saeed, Yasmeen, et al.
Veröffentlicht: (2026)
von: Saeed, Yasmeen, et al.
Veröffentlicht: (2026)
Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks
von: Alshehhi, Maitha, et al.
Veröffentlicht: (2025)
von: Alshehhi, Maitha, et al.
Veröffentlicht: (2025)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
LLaVA-SLT: Visual Language Tuning for Sign Language Translation
von: Liang, Han, et al.
Veröffentlicht: (2024)
von: Liang, Han, et al.
Veröffentlicht: (2024)
Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
von: Aboeitta, Ahmed, et al.
Veröffentlicht: (2025)
von: Aboeitta, Ahmed, et al.
Veröffentlicht: (2025)
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
Delta-LLaVA: Base-then-Specialize Alignment for Token-Efficient Vision-Language Models
von: Zamini, Mohamad, et al.
Veröffentlicht: (2025)
von: Zamini, Mohamad, et al.
Veröffentlicht: (2025)
LLaVA-LE: Large Language-and-Vision Assistant for Lunar Exploration
von: Inal, Gokce, et al.
Veröffentlicht: (2026)
von: Inal, Gokce, et al.
Veröffentlicht: (2026)
Dr-LLaVA: Visual Instruction Tuning with Symbolic Clinical Grounding
von: Sun, Shenghuan, et al.
Veröffentlicht: (2024)
von: Sun, Shenghuan, et al.
Veröffentlicht: (2024)
Single-Temporal Supervised Learning for Universal Remote Sensing Change Detection
von: Zheng, Zhuo, et al.
Veröffentlicht: (2024)
von: Zheng, Zhuo, et al.
Veröffentlicht: (2024)
Efficient Remote Sensing Change Detection with Change State Space Models
von: Ghazaei, Elman, et al.
Veröffentlicht: (2025)
von: Ghazaei, Elman, et al.
Veröffentlicht: (2025)
Yo'LLaVA: Your Personalized Language and Vision Assistant
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
GeoVLM-R1: Reinforcement Fine-Tuning for Improved Remote Sensing Reasoning
von: Fiaz, Mustansar, et al.
Veröffentlicht: (2025)
von: Fiaz, Mustansar, et al.
Veröffentlicht: (2025)
EfficientLLaVA:Generalizable Auto-Pruning for Large Vision-language Models
von: Liang, Yinan, et al.
Veröffentlicht: (2025)
von: Liang, Yinan, et al.
Veröffentlicht: (2025)
SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding
von: Luo, Junwei, et al.
Veröffentlicht: (2024)
von: Luo, Junwei, et al.
Veröffentlicht: (2024)
Remote Sensing Change Detection via Weak Temporal Supervision
von: Bou, Xavier, et al.
Veröffentlicht: (2026)
von: Bou, Xavier, et al.
Veröffentlicht: (2026)
ChestGPT: Integrating Large Language Models and Vision Transformers for Disease Detection and Localization in Chest X-Rays
von: Khan, Shehroz S., et al.
Veröffentlicht: (2025)
von: Khan, Shehroz S., et al.
Veröffentlicht: (2025)
LLaVA-c: Continual Improved Visual Instruction Tuning
von: Liu, Wenzhuo, et al.
Veröffentlicht: (2025)
von: Liu, Wenzhuo, et al.
Veröffentlicht: (2025)
Leveraging Fine-Grained Information and Noise Decoupling for Remote Sensing Change Detection
von: Du, Qiangang, et al.
Veröffentlicht: (2024)
von: Du, Qiangang, et al.
Veröffentlicht: (2024)
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
von: Jahagirdar, Soumya, et al.
Veröffentlicht: (2026)
von: Jahagirdar, Soumya, et al.
Veröffentlicht: (2026)
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
von: Lu, Weiheng, et al.
Veröffentlicht: (2024)
von: Lu, Weiheng, et al.
Veröffentlicht: (2024)
ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
von: Ye, Xubing, et al.
Veröffentlicht: (2024)
von: Ye, Xubing, et al.
Veröffentlicht: (2024)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
von: Lin, Bin, et al.
Veröffentlicht: (2024)
von: Lin, Bin, et al.
Veröffentlicht: (2024)
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
von: Xu, Guowei, et al.
Veröffentlicht: (2024)
von: Xu, Guowei, et al.
Veröffentlicht: (2024)
PE-CLIP: A Parameter-Efficient Fine-Tuning of Vision Language Models for Dynamic Facial Expression Recognition
von: Saadi, Ibtissam, et al.
Veröffentlicht: (2025)
von: Saadi, Ibtissam, et al.
Veröffentlicht: (2025)
Joint Spatio-Temporal Modeling for the Semantic Change Detection in Remote Sensing Images
von: Ding, Lei, et al.
Veröffentlicht: (2022)
von: Ding, Lei, et al.
Veröffentlicht: (2022)
NeXt2Former-CD: Efficient Remote Sensing Change Detection with Modern Vision Architectures
von: Wang, Yufan, et al.
Veröffentlicht: (2026)
von: Wang, Yufan, et al.
Veröffentlicht: (2026)
GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing
von: Ma, Xianzhi, et al.
Veröffentlicht: (2025)
von: Ma, Xianzhi, et al.
Veröffentlicht: (2025)
CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging
von: Imam, Raza, et al.
Veröffentlicht: (2024)
von: Imam, Raza, et al.
Veröffentlicht: (2024)
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression
von: Zhu, Yuke, et al.
Veröffentlicht: (2024)
von: Zhu, Yuke, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ChatENV: An Interactive Vision-Language Model for Sensor-Guided Environmental Monitoring and Scenario Simulation
von: Elgendy, Hosam, et al.
Veröffentlicht: (2025) -
Not Only Grey Matter: OmniBrain for Robust Multimodal Classification of Alzheimer's Disease
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2025) -
PulmoFusion: Advancing Pulmonary Health with Efficient Multi-Modal Fusion
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2025) -
GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
von: Wang, Fengxiang, et al.
Veröffentlicht: (2025) -
Vision-Language Models for Edge Networks: A Comprehensive Survey
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2025)