Lightweight Multimodal Adaptation of Vision Language Models for Species Recognition and Habitat Context Interpretation in Drone Thermal Imagery
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Hao, Qiu, Fang, Dong, Fangchao, Yang, Defei, Bohnett, Eve, An, Li |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-perspective monitoring of wildlife and human activities from camera traps and drones with deep learning models
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
AMMKD: Adaptive Multimodal Multi-teacher Distillation for Lightweight Vision-Language Models
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
von: Peng, Jingwei, et al.
Veröffentlicht: (2025)
von: Peng, Jingwei, et al.
Veröffentlicht: (2025)
Multi-Species Object Detection in Drone Imagery for Population Monitoring of Endangered Animals
von: Sankaran, Sowmya
Veröffentlicht: (2024)
von: Sankaran, Sowmya
Veröffentlicht: (2024)
Towards Multimodal In-Context Learning for Vision & Language Models
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
DroneSplat: 3D Gaussian Splatting for Robust 3D Reconstruction from In-the-Wild Drone Imagery
von: Tang, Jiadong, et al.
Veröffentlicht: (2025)
von: Tang, Jiadong, et al.
Veröffentlicht: (2025)
Understanding Representation Gaps Across Scales in Tropical Tree Species Classification from Drone Imagery
von: Saha, Sulagna, et al.
Veröffentlicht: (2026)
von: Saha, Sulagna, et al.
Veröffentlicht: (2026)
DroneScan-YOLO: Redundancy-Aware Lightweight Detection for Tiny Objects in UAV Imagery
von: Bellec, Yann V.
Veröffentlicht: (2026)
von: Bellec, Yann V.
Veröffentlicht: (2026)
Lightweight Transformer-Driven Segmentation of Hotspots and Snail Trails in Solar PV Thermal Imagery
von: Joshi, Deepak, et al.
Veröffentlicht: (2025)
von: Joshi, Deepak, et al.
Veröffentlicht: (2025)
Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning
von: Liu, Shih-Wen, et al.
Veröffentlicht: (2025)
von: Liu, Shih-Wen, et al.
Veröffentlicht: (2025)
Multimodal Protein Language Models for Enzyme Kinetic Parameters: From Substrate Recognition to Conformational Adaptation
von: Wang, Fei, et al.
Veröffentlicht: (2026)
von: Wang, Fei, et al.
Veröffentlicht: (2026)
ThermEval: A Structured Benchmark for Evaluation of Vision-Language Models on Thermal Imagery
von: Shrivastava, Ayush, et al.
Veröffentlicht: (2026)
von: Shrivastava, Ayush, et al.
Veröffentlicht: (2026)
Multimodal Interpretation of Remote Sensing Images: Dynamic Resolution Input Strategy and Multi-scale Vision-Language Alignment Mechanism
von: Zhang, Siyu, et al.
Veröffentlicht: (2025)
von: Zhang, Siyu, et al.
Veröffentlicht: (2025)
IIR-VLM: In-Context Instance-level Recognition for Large Vision-Language Models
von: Shi, Liang, et al.
Veröffentlicht: (2026)
von: Shi, Liang, et al.
Veröffentlicht: (2026)
Anthropogenic Regional Adaptation in Multimodal Vision-Language Model
von: Cahyawijaya, Samuel, et al.
Veröffentlicht: (2026)
von: Cahyawijaya, Samuel, et al.
Veröffentlicht: (2026)
Lightweight Unsupervised Federated Learning with Pretrained Vision Language Model
von: Yan, Hao, et al.
Veröffentlicht: (2024)
von: Yan, Hao, et al.
Veröffentlicht: (2024)
Leveraging YOLO-World and GPT-4V LMMs for Zero-Shot Person Detection and Action Recognition in Drone Imagery
von: Limberg, Christian, et al.
Veröffentlicht: (2024)
von: Limberg, Christian, et al.
Veröffentlicht: (2024)
Saliency-Guided Deep Learning for Bridge Defect Detection in Drone Imagery
von: Hebbache, Loucif, et al.
Veröffentlicht: (2025)
von: Hebbache, Loucif, et al.
Veröffentlicht: (2025)
SAGA: Semantic-Aware Gray color Augmentation for Visible-to-Thermal Domain Adaptation across Multi-View Drone and Ground-Based Vision Systems
von: D, Manjunath, et al.
Veröffentlicht: (2025)
von: D, Manjunath, et al.
Veröffentlicht: (2025)
Machine Vision Therapy: Multimodal Large Language Models Can Enhance Visual Robustness via Denoising In-Context Learning
von: Huang, Zhuo, et al.
Veröffentlicht: (2023)
von: Huang, Zhuo, et al.
Veröffentlicht: (2023)
CROPE: Evaluating In-Context Adaptation of Vision and Language Models to Culture-Specific Concepts
von: Nikandrou, Malvina, et al.
Veröffentlicht: (2024)
von: Nikandrou, Malvina, et al.
Veröffentlicht: (2024)
LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation
von: Li, Zhenshi, et al.
Veröffentlicht: (2024)
von: Li, Zhenshi, et al.
Veröffentlicht: (2024)
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks
von: Hu, Yuanze, et al.
Veröffentlicht: (2025)
von: Hu, Yuanze, et al.
Veröffentlicht: (2025)
Recov-Vision: Linking Street View Imagery and Vision-Language Models for Post-Disaster Recovery
von: Xiao, Yiming, et al.
Veröffentlicht: (2025)
von: Xiao, Yiming, et al.
Veröffentlicht: (2025)
SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning
von: Wu, Xue, et al.
Veröffentlicht: (2026)
von: Wu, Xue, et al.
Veröffentlicht: (2026)
Cascade Prompt Learning for Vision-Language Model Adaptation
von: Wu, Ge, et al.
Veröffentlicht: (2024)
von: Wu, Ge, et al.
Veröffentlicht: (2024)
Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models
von: Khan, Sumra, et al.
Veröffentlicht: (2026)
von: Khan, Sumra, et al.
Veröffentlicht: (2026)
Image Recognition with Online Lightweight Vision Transformer: A Survey
von: Zhang, Zherui, et al.
Veröffentlicht: (2025)
von: Zhang, Zherui, et al.
Veröffentlicht: (2025)
Assessing Building Heat Resilience Using UAV and Street-View Imagery with Coupled Global Context Vision Transformer
von: Knoblauch, Steffen, et al.
Veröffentlicht: (2026)
von: Knoblauch, Steffen, et al.
Veröffentlicht: (2026)
LLM-Powered Flood Depth Estimation from Social Media Imagery: A Vision-Language Model Framework with Mechanistic Interpretability for Transportation Resilience
von: Fuad, Nafis, et al.
Veröffentlicht: (2026)
von: Fuad, Nafis, et al.
Veröffentlicht: (2026)
Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
von: Lee, Phillip Y., et al.
Veröffentlicht: (2025)
von: Lee, Phillip Y., et al.
Veröffentlicht: (2025)
Vision-Enhanced Large Language Models for High-Resolution Image Synthesis and Multimodal Data Interpretation
von: KV, Karthikeya
Veröffentlicht: (2025)
von: KV, Karthikeya
Veröffentlicht: (2025)
Dynamic Multimodal Prototype Learning in Vision-Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025)
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
von: Stan, Gabriela Ben Melech, et al.
Veröffentlicht: (2024)
von: Stan, Gabriela Ben Melech, et al.
Veröffentlicht: (2024)
Evolving Prompt Adaptation for Vision-Language Models
von: Zhang, Enming, et al.
Veröffentlicht: (2026)
von: Zhang, Enming, et al.
Veröffentlicht: (2026)
Fair Context Learning for Evidence-Balanced Test-Time Adaptation in Vision-Language Models
von: Yun, Sanggeon, et al.
Veröffentlicht: (2026)
von: Yun, Sanggeon, et al.
Veröffentlicht: (2026)
Interpretable Bilingual Multimodal Large Language Model for Diverse Biomedical Tasks
von: Wang, Lehan, et al.
Veröffentlicht: (2024)
von: Wang, Lehan, et al.
Veröffentlicht: (2024)
CLIPSwarm: Generating Drone Shows from Text Prompts with Vision-Language Models
von: Pueyo, Pablo, et al.
Veröffentlicht: (2024)
von: Pueyo, Pablo, et al.
Veröffentlicht: (2024)
VLN-Pilot: Large Vision-Language Model as an Autonomous Indoor Drone Operator
von: Dominguez-Dager, Bessie, et al.
Veröffentlicht: (2026)
von: Dominguez-Dager, Bessie, et al.
Veröffentlicht: (2026)
Dynamic Rank Adaptation for Vision-Language Models
von: Wang, Jiahui, et al.
Veröffentlicht: (2025)
von: Wang, Jiahui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multi-perspective monitoring of wildlife and human activities from camera traps and drones with deep learning models
von: Chen, Hao, et al.
Veröffentlicht: (2025) -
AMMKD: Adaptive Multimodal Multi-teacher Distillation for Lightweight Vision-Language Models
von: Li, Yuqi, et al.
Veröffentlicht: (2025) -
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
von: Peng, Jingwei, et al.
Veröffentlicht: (2025) -
Multi-Species Object Detection in Drone Imagery for Population Monitoring of Endangered Animals
von: Sankaran, Sowmya
Veröffentlicht: (2024) -
Towards Multimodal In-Context Learning for Vision & Language Models
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)