bi-modal textual prompt learning for vision-language models in remote sensing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kashyap, Pankhi, Singha, Mainak, Banerjee, Biplab |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AD-CLIP: Adapting Domains in Prompt Space Using CLIP
von: Singha, Mainak, et al.
Veröffentlicht: (2023)
von: Singha, Mainak, et al.
Veröffentlicht: (2023)
Elevating All Zero-Shot Sketch-Based Image Retrieval Through Multimodal Prompt Learning
von: Singha, Mainak, et al.
Veröffentlicht: (2024)
von: Singha, Mainak, et al.
Veröffentlicht: (2024)
MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP
von: Chaudhary, Aditya, et al.
Veröffentlicht: (2026)
von: Chaudhary, Aditya, et al.
Veröffentlicht: (2026)
Unknown Prompt, the only Lacuna: Unveiling CLIP's Potential for Open Domain Generalization
von: Singha, Mainak, et al.
Veröffentlicht: (2024)
von: Singha, Mainak, et al.
Veröffentlicht: (2024)
GraphVL: Graph-Enhanced Semantic Modeling via Vision-Language Models for Generalized Class Discovery
von: Solanki, Bhupendra, et al.
Veröffentlicht: (2024)
von: Solanki, Bhupendra, et al.
Veröffentlicht: (2024)
BioVLM: Routing Prompts, Not Parameters, for Cross-Modality Generalization in Biomedical VLMs
von: Singha, Mainak, et al.
Veröffentlicht: (2026)
von: Singha, Mainak, et al.
Veröffentlicht: (2026)
COSMo: CLIP Talks on Open-Set Multi-Target Domain Adaptation
von: Monga, Munish, et al.
Veröffentlicht: (2024)
von: Monga, Munish, et al.
Veröffentlicht: (2024)
FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models
von: Singha, Mainak, et al.
Veröffentlicht: (2025)
von: Singha, Mainak, et al.
Veröffentlicht: (2025)
CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation
von: Singha, Mainak, et al.
Veröffentlicht: (2026)
von: Singha, Mainak, et al.
Veröffentlicht: (2026)
Reconstruction Guided Few-shot Network For Remote Sensing Image Classification
von: Jaiswal, Mohit, et al.
Veröffentlicht: (2026)
von: Jaiswal, Mohit, et al.
Veröffentlicht: (2026)
SDHSI-Net: Learning Better Representations for Hyperspectral Images via Self-Distillation
von: Singh, Prachet Dev, et al.
Veröffentlicht: (2026)
von: Singh, Prachet Dev, et al.
Veröffentlicht: (2026)
FrogDogNet: Fourier frequency Retained visual prompt Output Guidance for Domain Generalization of CLIP in Remote Sensing
von: Gunduboina, Hariseetharam, et al.
Veröffentlicht: (2025)
von: Gunduboina, Hariseetharam, et al.
Veröffentlicht: (2025)
On the test-time zero-shot generalization of vision-language models: Do we really need prompt learning?
von: Zanella, Maxime, et al.
Veröffentlicht: (2024)
von: Zanella, Maxime, et al.
Veröffentlicht: (2024)
DepthSeg: Depth prompting in remote sensing semantic segmentation
von: Zhou, Ning, et al.
Veröffentlicht: (2025)
von: Zhou, Ning, et al.
Veröffentlicht: (2025)
OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIP
von: C, Mohamad Hassan N, et al.
Veröffentlicht: (2025)
von: C, Mohamad Hassan N, et al.
Veröffentlicht: (2025)
Noise-aware few-shot learning through bi-directional multi-view prompt alignment
von: Niu, Lu, et al.
Veröffentlicht: (2026)
von: Niu, Lu, et al.
Veröffentlicht: (2026)
CDAD-Net: Bridging Domain Gaps in Generalized Category Discovery
von: Rongali, Sai Bhargav, et al.
Veröffentlicht: (2024)
von: Rongali, Sai Bhargav, et al.
Veröffentlicht: (2024)
FALCON: Few-Shot Adversarial Learning for Cross-Domain Medical Image Segmentation
von: Fayjie, Abdur R., et al.
Veröffentlicht: (2026)
von: Fayjie, Abdur R., et al.
Veröffentlicht: (2026)
Deep learning-based interactive segmentation in remote sensing
von: Wang, Zhe, et al.
Veröffentlicht: (2023)
von: Wang, Zhe, et al.
Veröffentlicht: (2023)
A multi-modal vision-language model for generalizable annotation-free pathology localization
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
CromSS: Cross-modal pre-training with noisy labels for remote sensing image segmentation
von: Liu, Chenying, et al.
Veröffentlicht: (2024)
von: Liu, Chenying, et al.
Veröffentlicht: (2024)
Leveraging feature communication in federated learning for remote sensing image classification
von: Duong, Anh-Kiet, et al.
Veröffentlicht: (2024)
von: Duong, Anh-Kiet, et al.
Veröffentlicht: (2024)
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
von: Nassar, Ahmed, et al.
Veröffentlicht: (2025)
von: Nassar, Ahmed, et al.
Veröffentlicht: (2025)
Zero-shot large vision-language model prompting for automated bone identification in paleoradiology x-ray archives
von: Dong, Owen, et al.
Veröffentlicht: (2026)
von: Dong, Owen, et al.
Veröffentlicht: (2026)
AutoOEP -- A Multi-modal Framework for Online Exam Proctoring
von: Naveen, Aryan Kashyap
Veröffentlicht: (2025)
von: Naveen, Aryan Kashyap
Veröffentlicht: (2025)
The in-context inductive biases of vision-language models differ across modalities
von: Allen, Kelsey, et al.
Veröffentlicht: (2025)
von: Allen, Kelsey, et al.
Veröffentlicht: (2025)
Unified modality separation: A vision-language framework for unsupervised domain adaptation
von: Li, Xinyao, et al.
Veröffentlicht: (2025)
von: Li, Xinyao, et al.
Veröffentlicht: (2025)
GeoMeld: Toward Semantically Grounded Foundation Models for Remote Sensing
von: Hasan, Maram, et al.
Veröffentlicht: (2026)
von: Hasan, Maram, et al.
Veröffentlicht: (2026)
VLA-Mark: A cross modal watermark for large vision-language alignment model
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
Cross-modal linkage risk in clinical vision-language models
von: Arasteh, Soroosh Tayebi, et al.
Veröffentlicht: (2026)
von: Arasteh, Soroosh Tayebi, et al.
Veröffentlicht: (2026)
Extending global-local view alignment for self-supervised learning with remote sensing imagery
von: Wanyan, Xinye, et al.
Veröffentlicht: (2023)
von: Wanyan, Xinye, et al.
Veröffentlicht: (2023)
Enhancing medical vision-language contrastive learning via inter-matching relation modelling
von: Li, Mingjian, et al.
Veröffentlicht: (2024)
von: Li, Mingjian, et al.
Veröffentlicht: (2024)
Local-Global Context-Aware and Structure-Preserving Image Super-Resolution
von: Palit, Sanchar, et al.
Veröffentlicht: (2025)
von: Palit, Sanchar, et al.
Veröffentlicht: (2025)
HIDISC: A Hyperbolic Framework for Domain Generalization with Generalized Category Discovery
von: Rathore, Vaibhav, et al.
Veröffentlicht: (2025)
von: Rathore, Vaibhav, et al.
Veröffentlicht: (2025)
Label-Efficient Hyperspectral Image Classification via Spectral FiLM Modulation of Low-Level Pretrained Diffusion Features
von: Hu, Yuzhen, et al.
Veröffentlicht: (2025)
von: Hu, Yuzhen, et al.
Veröffentlicht: (2025)
Oriented object detection in optical remote sensing images using deep learning: a survey
von: Wang, Kun, et al.
Veröffentlicht: (2023)
von: Wang, Kun, et al.
Veröffentlicht: (2023)
Leveraging knowledge distillation for partial multi-task learning from multiple remote sensing datasets
von: Lê, Hoàng-Ân, et al.
Veröffentlicht: (2024)
von: Lê, Hoàng-Ân, et al.
Veröffentlicht: (2024)
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
von: Xing, Yang, et al.
Veröffentlicht: (2026)
von: Xing, Yang, et al.
Veröffentlicht: (2026)
An analysis of vision-language models for fabric retrieval
von: Giuliari, Francesco, et al.
Veröffentlicht: (2025)
von: Giuliari, Francesco, et al.
Veröffentlicht: (2025)
A generalised pre-training strategy for deep learning networks in semantic segmentation of remotely sensed images
von: Fang, Yuan, et al.
Veröffentlicht: (2026)
von: Fang, Yuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
AD-CLIP: Adapting Domains in Prompt Space Using CLIP
von: Singha, Mainak, et al.
Veröffentlicht: (2023) -
Elevating All Zero-Shot Sketch-Based Image Retrieval Through Multimodal Prompt Learning
von: Singha, Mainak, et al.
Veröffentlicht: (2024) -
MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP
von: Chaudhary, Aditya, et al.
Veröffentlicht: (2026) -
Unknown Prompt, the only Lacuna: Unveiling CLIP's Potential for Open Domain Generalization
von: Singha, Mainak, et al.
Veröffentlicht: (2024) -
GraphVL: Graph-Enhanced Semantic Modeling via Vision-Language Models for Generalized Class Discovery
von: Solanki, Bhupendra, et al.
Veröffentlicht: (2024)