BioVLM: Routing Prompts, Not Parameters, for Cross-Modality Generalization in Biomedical VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Singha, Mainak, Gupta, Tanisha, Jha, Ankit, Khan, Muhammad Haris, Ghosh, Sayantani, Banerjee, Biplab |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AD-CLIP: Adapting Domains in Prompt Space Using CLIP
by: Singha, Mainak, et al.
Published: (2023)
by: Singha, Mainak, et al.
Published: (2023)
OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIP
by: C, Mohamad Hassan N, et al.
Published: (2025)
by: C, Mohamad Hassan N, et al.
Published: (2025)
Elevating All Zero-Shot Sketch-Based Image Retrieval Through Multimodal Prompt Learning
by: Singha, Mainak, et al.
Published: (2024)
by: Singha, Mainak, et al.
Published: (2024)
MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP
by: Chaudhary, Aditya, et al.
Published: (2026)
by: Chaudhary, Aditya, et al.
Published: (2026)
Unknown Prompt, the only Lacuna: Unveiling CLIP's Potential for Open Domain Generalization
by: Singha, Mainak, et al.
Published: (2024)
by: Singha, Mainak, et al.
Published: (2024)
CDAD-Net: Bridging Domain Gaps in Generalized Category Discovery
by: Rongali, Sai Bhargav, et al.
Published: (2024)
by: Rongali, Sai Bhargav, et al.
Published: (2024)
GraphVL: Graph-Enhanced Semantic Modeling via Vision-Language Models for Generalized Class Discovery
by: Solanki, Bhupendra, et al.
Published: (2024)
by: Solanki, Bhupendra, et al.
Published: (2024)
FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models
by: Singha, Mainak, et al.
Published: (2025)
by: Singha, Mainak, et al.
Published: (2025)
bi-modal textual prompt learning for vision-language models in remote sensing
by: Kashyap, Pankhi, et al.
Published: (2026)
by: Kashyap, Pankhi, et al.
Published: (2026)
COSMo: CLIP Talks on Open-Set Multi-Target Domain Adaptation
by: Monga, Munish, et al.
Published: (2024)
by: Monga, Munish, et al.
Published: (2024)
FrogDogNet: Fourier frequency Retained visual prompt Output Guidance for Domain Generalization of CLIP in Remote Sensing
by: Gunduboina, Hariseetharam, et al.
Published: (2025)
by: Gunduboina, Hariseetharam, et al.
Published: (2025)
Reconstruction Guided Few-shot Network For Remote Sensing Image Classification
by: Jaiswal, Mohit, et al.
Published: (2026)
by: Jaiswal, Mohit, et al.
Published: (2026)
SDHSI-Net: Learning Better Representations for Hyperspectral Images via Self-Distillation
by: Singh, Prachet Dev, et al.
Published: (2026)
by: Singh, Prachet Dev, et al.
Published: (2026)
GeoMeld: Toward Semantically Grounded Foundation Models for Remote Sensing
by: Hasan, Maram, et al.
Published: (2026)
by: Hasan, Maram, et al.
Published: (2026)
CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation
by: Singha, Mainak, et al.
Published: (2026)
by: Singha, Mainak, et al.
Published: (2026)
In the Era of Prompt Learning with Vision-Language Models
by: Jha, Ankit
Published: (2024)
by: Jha, Ankit
Published: (2024)
HIDISC: A Hyperbolic Framework for Domain Generalization with Generalized Category Discovery
by: Rathore, Vaibhav, et al.
Published: (2025)
by: Rathore, Vaibhav, et al.
Published: (2025)
AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection
by: Jiang, Yichen, et al.
Published: (2025)
by: Jiang, Yichen, et al.
Published: (2025)
Federated Cross-Modal Style-Aware Prompt Generation
by: Prasad, Suraj, et al.
Published: (2025)
by: Prasad, Suraj, et al.
Published: (2025)
CR-JEPA: Cross-Modal Joint-Embedding Predictive Learning for Remote Sensing Image Retrieval
by: Hossain, Md Aminur, et al.
Published: (2026)
by: Hossain, Md Aminur, et al.
Published: (2026)
Discovery of a 13-Sharpe OOS Factor: Drift Regimes Unlock Hidden Cross-Sectional Predictability
by: Singha, Mainak
Published: (2025)
by: Singha, Mainak
Published: (2025)
O-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language Models
by: Sharifdeen, Ashshak, et al.
Published: (2025)
by: Sharifdeen, Ashshak, et al.
Published: (2025)
MedObvious: Exposing the Medical Moravec's Paradox in VLMs via Clinical Triage
by: Khan, Ufaq, et al.
Published: (2026)
by: Khan, Ufaq, et al.
Published: (2026)
Divergent Domains, Convergent Grading: Enhancing Generalization in Diabetic Retinopathy Grading
by: Chokuwa, Sharon, et al.
Published: (2024)
by: Chokuwa, Sharon, et al.
Published: (2024)
HQ-JEPA: Hybrid Quantum Joint-Embedding Predictive Architecture for Cross-Modal Remote Sensing Representation Learning
by: Hossain, Md Aminur, et al.
Published: (2026)
by: Hossain, Md Aminur, et al.
Published: (2026)
Improving Pseudo-labelling and Enhancing Robustness for Semi-Supervised Domain Generalization
by: Khan, Adnan, et al.
Published: (2024)
by: Khan, Adnan, et al.
Published: (2024)
Robust and Label-Efficient Deep Waste Detection
by: Abid, Hassan, et al.
Published: (2025)
by: Abid, Hassan, et al.
Published: (2025)
Revised Regularization for Efficient Continual Learning through Correlation-Based Parameter Update in Bayesian Neural Networks
by: Palit, Sanchar, et al.
Published: (2024)
by: Palit, Sanchar, et al.
Published: (2024)
Foundation Models and Adaptive Feature Selection: A Synergistic Approach to Video Question Answering
by: Rongali, Sai Bhargav, et al.
Published: (2024)
by: Rongali, Sai Bhargav, et al.
Published: (2024)
vMFCoOp: Towards Equilibrium on a Unified Hyperspherical Manifold for Prompting Biomedical VLMs
by: Shao, Minye, et al.
Published: (2025)
by: Shao, Minye, et al.
Published: (2025)
CountZES: Counting via Zero-Shot Exemplar Selection
by: Siddiqui, Muhammad Ibraheem, et al.
Published: (2025)
by: Siddiqui, Muhammad Ibraheem, et al.
Published: (2025)
MyVLM: Personalizing VLMs for User-Specific Queries
by: Alaluf, Yuval, et al.
Published: (2024)
by: Alaluf, Yuval, et al.
Published: (2024)
Pixels Don't Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning Supervision
by: Kuckreja, Kartik, et al.
Published: (2026)
by: Kuckreja, Kartik, et al.
Published: (2026)
Calibration-Aware Prompt Learning for Medical Vision-Language Models
by: Basu, Abhishek, et al.
Published: (2025)
by: Basu, Abhishek, et al.
Published: (2025)
BMIP: Bi-directional Modality Interaction Prompt Learning for VLM
by: Lv, Song-Lin, et al.
Published: (2025)
by: Lv, Song-Lin, et al.
Published: (2025)
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
by: Deria, Ankan, et al.
Published: (2025)
by: Deria, Ankan, et al.
Published: (2025)
Noise-Tolerant Few-Shot Unsupervised Adapter for Vision-Language Models
by: Ali, Eman, et al.
Published: (2023)
by: Ali, Eman, et al.
Published: (2023)
Multi-modal Generation via Cross-Modal In-Context Learning
by: Kumar, Amandeep, et al.
Published: (2024)
by: Kumar, Amandeep, et al.
Published: (2024)
Modality Invariant Multimodal Learning to Handle Missing Modalities: A Single-Branch Approach
by: Saeed, Muhammad Saad, et al.
Published: (2024)
by: Saeed, Muhammad Saad, et al.
Published: (2024)
Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos
by: Ding, Bonan, et al.
Published: (2026)
by: Ding, Bonan, et al.
Published: (2026)
Similar Items
-
AD-CLIP: Adapting Domains in Prompt Space Using CLIP
by: Singha, Mainak, et al.
Published: (2023) -
OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIP
by: C, Mohamad Hassan N, et al.
Published: (2025) -
Elevating All Zero-Shot Sketch-Based Image Retrieval Through Multimodal Prompt Learning
by: Singha, Mainak, et al.
Published: (2024) -
MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP
by: Chaudhary, Aditya, et al.
Published: (2026) -
Unknown Prompt, the only Lacuna: Unveiling CLIP's Potential for Open Domain Generalization
by: Singha, Mainak, et al.
Published: (2024)