Can Better Text Semantics in Prompt Tuning Improve VLM Generalization?
Fuente:
arXiv
Saved in:
| Main Authors: | Kuchibhotla, Hari Chandana, Kancheti, Sai Srinivas, Reddy, Abbavaram Gowtham, Balasubramanian, Vineeth N |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs
by: Sinha, Rohit, et al.
Published: (2026)
by: Sinha, Rohit, et al.
Published: (2026)
$\oslash$ Source Models Leak What They Shouldn't $\nrightarrow$: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimization
by: Devalapally, Arnav, et al.
Published: (2026)
by: Devalapally, Arnav, et al.
Published: (2026)
Detecting and Measuring Confounding Using Causal Mechanism Shifts
by: Reddy, Abbavaram Gowtham, et al.
Published: (2024)
by: Reddy, Abbavaram Gowtham, et al.
Published: (2024)
NESTER: An Adaptive Neurosymbolic Method for Causal Effect Estimation
by: Reddy, Abbavaram Gowtham, et al.
Published: (2022)
by: Reddy, Abbavaram Gowtham, et al.
Published: (2022)
Source-Free Domain Adaptation by Optimizing Batch-Wise Cosine Similarity
by: Pathak, Harsharaj, et al.
Published: (2026)
by: Pathak, Harsharaj, et al.
Published: (2026)
Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach
by: Khindkar, Vaishnavi, et al.
Published: (2024)
by: Khindkar, Vaishnavi, et al.
Published: (2024)
POET: Prompt Offset Tuning for Continual Human Action Adaptation
by: Garg, Prachi, et al.
Published: (2025)
by: Garg, Prachi, et al.
Published: (2025)
iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
by: Mehrotra, Sarthak, et al.
Published: (2025)
by: Mehrotra, Sarthak, et al.
Published: (2025)
C2FDrone: Coarse-to-Fine Drone-to-Drone Detection using Vision Transformer Networks
by: Rebbapragada, Sairam VC, et al.
Published: (2024)
by: Rebbapragada, Sairam VC, et al.
Published: (2024)
Advancing Ante-Hoc Explainable Models through Generative Adversarial Networks
by: Garg, Tanmay, et al.
Published: (2024)
by: Garg, Tanmay, et al.
Published: (2024)
Precise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class Imbalance
by: Santra, Sanchayan, et al.
Published: (2025)
by: Santra, Sanchayan, et al.
Published: (2025)
PromptSafe: Gated Prompt Tuning for Safe Text-to-Image Generation
by: Jing, Zonglei, et al.
Published: (2025)
by: Jing, Zonglei, et al.
Published: (2025)
Annotation-Free Class-Incremental Learning
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
LogicCBMs: Logic-Enhanced Concept-Based Learning
by: Vemuri, Deepika SN, et al.
Published: (2025)
by: Vemuri, Deepika SN, et al.
Published: (2025)
Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media
by: M, Megha Mariam K., et al.
Published: (2026)
by: M, Megha Mariam K., et al.
Published: (2026)
Mitigate One, Skew Another? Tackling Intersectional Biases in Text-to-Image Models
by: Shukla, Pushkar, et al.
Published: (2025)
by: Shukla, Pushkar, et al.
Published: (2025)
Open-Set Object Detection By Aligning Known Class Representations
by: Sarkar, Hiran, et al.
Published: (2024)
by: Sarkar, Hiran, et al.
Published: (2024)
On Evaluation of Vision Datasets and Models using Human Competency Frameworks
by: Ramachandran, Rahul, et al.
Published: (2024)
by: Ramachandran, Rahul, et al.
Published: (2024)
Semantic Hierarchical Prompt Tuning for Parameter-Efficient Fine-Tuning
by: Zhu, Haowei, et al.
Published: (2024)
by: Zhu, Haowei, et al.
Published: (2024)
Understanding Task Transfer in Vision-Language Models
by: Sachdeva, Bhuvan, et al.
Published: (2025)
by: Sachdeva, Bhuvan, et al.
Published: (2025)
Poetry in Pixels: Prompt Tuning for Poem Image Generation via Diffusion Models
by: Jamil, Sofia, et al.
Published: (2025)
by: Jamil, Sofia, et al.
Published: (2025)
VLM-Guided Adaptive Negative Prompting for Creative Generation
by: Golan, Shelly, et al.
Published: (2025)
by: Golan, Shelly, et al.
Published: (2025)
GeoVLM-R1: Reinforcement Fine-Tuning for Improved Remote Sensing Reasoning
by: Fiaz, Mustansar, et al.
Published: (2025)
by: Fiaz, Mustansar, et al.
Published: (2025)
MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM
by: Chen, Tao, et al.
Published: (2025)
by: Chen, Tao, et al.
Published: (2025)
FRAP: Faithful and Realistic Text-to-Image Generation with Adaptive Prompt Weighting
by: Jiang, Liyao, et al.
Published: (2024)
by: Jiang, Liyao, et al.
Published: (2024)
The Describe-Then-Generate Bottleneck: How VLM Descriptions Alter Image Generation Outcomes
by: Kodathala, Sai Varun, et al.
Published: (2025)
by: Kodathala, Sai Varun, et al.
Published: (2025)
BioVLM: Routing Prompts, Not Parameters, for Cross-Modality Generalization in Biomedical VLMs
by: Singha, Mainak, et al.
Published: (2026)
by: Singha, Mainak, et al.
Published: (2026)
SemPT: Semantic Prompt Tuning for Vision-Language Models
by: Shi, Xiao, et al.
Published: (2025)
by: Shi, Xiao, et al.
Published: (2025)
Flexible Control of 3D CT Generation via Text and Semantically-Defined Segmentation Prompts
by: Dai, Weicheng, et al.
Published: (2026)
by: Dai, Weicheng, et al.
Published: (2026)
SoC: Semantic Orthogonal Calibration for Test-Time Prompt Tuning
by: Fillioux, Leo, et al.
Published: (2026)
by: Fillioux, Leo, et al.
Published: (2026)
Robust Prompt Tuning for Vision-Language Models with Mild Semantic Noise
by: Gao, Yansheng, et al.
Published: (2025)
by: Gao, Yansheng, et al.
Published: (2025)
DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers
by: Ren, Li, et al.
Published: (2025)
by: Ren, Li, et al.
Published: (2025)
Towards A Better Metric for Text-to-Video Generation
by: Wu, Jay Zhangjie, et al.
Published: (2024)
by: Wu, Jay Zhangjie, et al.
Published: (2024)
BiasConnect: Investigating Bias Interactions in Text-to-Image Models
by: Shukla, Pushkar, et al.
Published: (2025)
by: Shukla, Pushkar, et al.
Published: (2025)
DGL: Dynamic Global-Local Prompt Tuning for Text-Video Retrieval
by: Yang, Xiangpeng, et al.
Published: (2024)
by: Yang, Xiangpeng, et al.
Published: (2024)
Text as Any-Modality for Zero-Shot Classification by Consistent Prompt Tuning
by: Wu, Xiangyu, et al.
Published: (2025)
by: Wu, Xiangyu, et al.
Published: (2025)
TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
by: Stergiou, Alexandros
Published: (2025)
by: Stergiou, Alexandros
Published: (2025)
Similar Items
-
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025) -
Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization
by: Kancheti, Sai Srinivas, et al.
Published: (2026) -
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
by: Kancheti, Sai Srinivas, et al.
Published: (2026) -
Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs
by: Sinha, Rohit, et al.
Published: (2026) -
$\oslash$ Source Models Leak What They Shouldn't $\nrightarrow$: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimization
by: Devalapally, Arnav, et al.
Published: (2026)