Learning to Prompt with Text Only Supervision for Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khattak, Muhammad Uzair, Naeem, Muhammad Ferjad, Naseer, Muzammal, Van Gool, Luc, Tombari, Federico |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
von: Kuzucu, Selim, et al.
Veröffentlicht: (2025)
von: Kuzucu, Selim, et al.
Veröffentlicht: (2025)
Human Pose Descriptions and Subject-Focused Attention for Improved Zero-Shot Transfer in Human-Centric Classification Tasks
von: Khan, Muhammad Saif Ullah, et al.
Veröffentlicht: (2024)
von: Khan, Muhammad Saif Ullah, et al.
Veröffentlicht: (2024)
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization
von: Hassan, Jameel, et al.
Veröffentlicht: (2023)
von: Hassan, Jameel, et al.
Veröffentlicht: (2023)
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
von: Kuzucu, Selim, et al.
Veröffentlicht: (2026)
von: Kuzucu, Selim, et al.
Veröffentlicht: (2026)
Toward a Diffusion-Based Generalist for Dense Vision Tasks
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
PromptSmooth: Certifying Robustness of Medical Vision-Language Models via Prompt Learning
von: Hussein, Noor, et al.
Veröffentlicht: (2024)
von: Hussein, Noor, et al.
Veröffentlicht: (2024)
Know Your Neighbors: Improving Single-View Reconstruction via Spatial Vision-Language Reasoning
von: Li, Rui, et al.
Veröffentlicht: (2024)
von: Li, Rui, et al.
Veröffentlicht: (2024)
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
von: Kukleva, Anna, et al.
Veröffentlicht: (2025)
von: Kukleva, Anna, et al.
Veröffentlicht: (2025)
Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning
von: Watawana, Hasindri, et al.
Veröffentlicht: (2024)
von: Watawana, Hasindri, et al.
Veröffentlicht: (2024)
Promptception: How Sensitive Are Large Multimodal Models to Prompts?
von: Ismithdeen, Mohamed Insaf, et al.
Veröffentlicht: (2025)
von: Ismithdeen, Mohamed Insaf, et al.
Veröffentlicht: (2025)
BAPLe: Backdoor Attacks on Medical Foundational Models using Prompt Learning
von: Hanif, Asif, et al.
Veröffentlicht: (2024)
von: Hanif, Asif, et al.
Veröffentlicht: (2024)
Self-supervised Shape Completion via Involution and Implicit Correspondences
von: Liu, Mengya, et al.
Veröffentlicht: (2024)
von: Liu, Mengya, et al.
Veröffentlicht: (2024)
Hierarchical Self-Supervised Adversarial Training for Robust Vision Models in Histopathology
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2025)
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2025)
LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts
von: Gani, Hanan, et al.
Veröffentlicht: (2023)
von: Gani, Hanan, et al.
Veröffentlicht: (2023)
GiT: Towards Generalist Vision Transformer through Universal Language Interface
von: Wang, Haiyang, et al.
Veröffentlicht: (2024)
von: Wang, Haiyang, et al.
Veröffentlicht: (2024)
MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
von: Segu, Mattia, et al.
Veröffentlicht: (2025)
von: Segu, Mattia, et al.
Veröffentlicht: (2025)
Active Data Curation Effectively Distills Large-Scale Multimodal Models
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2024)
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2024)
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2024)
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2024)
One2Any: One-Reference 6D Pose Estimation for Any Object
von: Liu, Mengya, et al.
Veröffentlicht: (2025)
von: Liu, Mengya, et al.
Veröffentlicht: (2025)
XDT-CXR: Investigating Cross-Disease Transferability in Zero-Shot Binary Classification of Chest X-Rays
von: Rahman, Umaima, et al.
Veröffentlicht: (2024)
von: Rahman, Umaima, et al.
Veröffentlicht: (2024)
AgriCLIP: Adapting CLIP for Agriculture and Livestock via Domain-Specialized Cross-Model Alignment
von: Nawaz, Umair, et al.
Veröffentlicht: (2024)
von: Nawaz, Umair, et al.
Veröffentlicht: (2024)
Bayesian Self-Training for Semi-Supervised 3D Segmentation
von: Unal, Ozan, et al.
Veröffentlicht: (2024)
von: Unal, Ozan, et al.
Veröffentlicht: (2024)
InseRF: Text-Driven Generative Object Insertion in Neural 3D Scenes
von: Shahbazi, Mohamad, et al.
Veröffentlicht: (2024)
von: Shahbazi, Mohamad, et al.
Veröffentlicht: (2024)
Multi-modal Generation via Cross-Modal In-Context Learning
von: Kumar, Amandeep, et al.
Veröffentlicht: (2024)
von: Kumar, Amandeep, et al.
Veröffentlicht: (2024)
Y-CA-Net: A Convolutional Attention Based Network for Volumetric Medical Image Segmentation
von: Sharif, Muhammad Hamza, et al.
Veröffentlicht: (2024)
von: Sharif, Muhammad Hamza, et al.
Veröffentlicht: (2024)
Adversarial Appearance Learning in Augmented Cityscapes for Pedestrian Recognition in Autonomous Driving
von: Savkin, Artem, et al.
Veröffentlicht: (2025)
von: Savkin, Artem, et al.
Veröffentlicht: (2025)
STEREO: A Two-Stage Framework for Adversarially Robust Concept Erasing from Text-to-Image Diffusion Models
von: Srivatsan, Koushik, et al.
Veröffentlicht: (2024)
von: Srivatsan, Koushik, et al.
Veröffentlicht: (2024)
Investigating the Effectiveness of Cross-Attention to Unlock Zero-Shot Editing of Text-to-Video Diffusion Models
von: Motamed, Saman, et al.
Veröffentlicht: (2024)
von: Motamed, Saman, et al.
Veröffentlicht: (2024)
Splat-SLAM: Globally Optimized RGB-only SLAM with 3D Gaussians
von: Sandström, Erik, et al.
Veröffentlicht: (2024)
von: Sandström, Erik, et al.
Veröffentlicht: (2024)
Lost in Translation? Vocabulary Alignment for Source-Free Adaptation in Open-Vocabulary Semantic Segmentation
von: Mazzucco, Silvio, et al.
Veröffentlicht: (2025)
von: Mazzucco, Silvio, et al.
Veröffentlicht: (2025)
Language Guided Domain Generalized Medical Image Segmentation
von: Kunhimon, Shahina, et al.
Veröffentlicht: (2024)
von: Kunhimon, Shahina, et al.
Veröffentlicht: (2024)
Camera-Only 3D Panoptic Scene Completion for Autonomous Driving through Differentiable Object Shapes
von: Marinello, Nicola, et al.
Veröffentlicht: (2025)
von: Marinello, Nicola, et al.
Veröffentlicht: (2025)
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
von: Motamed, Saman, et al.
Veröffentlicht: (2023)
von: Motamed, Saman, et al.
Veröffentlicht: (2023)
Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models
von: Yi, Jinhui, et al.
Veröffentlicht: (2024)
von: Yi, Jinhui, et al.
Veröffentlicht: (2024)
Calibration-Aware Prompt Learning for Medical Vision-Language Models
von: Basu, Abhishek, et al.
Veröffentlicht: (2025)
von: Basu, Abhishek, et al.
Veröffentlicht: (2025)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
von: Dharmasiri, Amaya, et al.
Veröffentlicht: (2024)
von: Dharmasiri, Amaya, et al.
Veröffentlicht: (2024)
Self-Supervised Learning with a Multi-Task Latent Space Objective
von: De Plaen, Pierre-François, et al.
Veröffentlicht: (2026)
von: De Plaen, Pierre-François, et al.
Veröffentlicht: (2026)
Incremental Object Detection with Prompt-based Methods
von: Neuwirth-Trapp, Matthias, et al.
Veröffentlicht: (2025)
von: Neuwirth-Trapp, Matthias, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024) -
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
von: Kuzucu, Selim, et al.
Veröffentlicht: (2025) -
Human Pose Descriptions and Subject-Focused Attention for Improved Zero-Shot Transfer in Human-Centric Classification Tasks
von: Khan, Muhammad Saif Ullah, et al.
Veröffentlicht: (2024) -
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024) -
Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization
von: Hassan, Jameel, et al.
Veröffentlicht: (2023)