Prompting Large Vision-Language Models for Compositional Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ossowski, Timothy, Jiang, Ming, Hu, Junjie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OLIVE: Object Level In-Context Visual Embeddings
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024)
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024)
Physical Prompt Injection Attacks on Large Vision-Language Models
von: Ling, Chen, et al.
Veröffentlicht: (2026)
von: Ling, Chen, et al.
Veröffentlicht: (2026)
Attention Prompting on Image for Large Vision-Language Models
von: Yu, Runpeng, et al.
Veröffentlicht: (2024)
von: Yu, Runpeng, et al.
Veröffentlicht: (2024)
Adversarial Prompt Distillation for Vision-Language Models
von: Luo, Lin, et al.
Veröffentlicht: (2024)
von: Luo, Lin, et al.
Veröffentlicht: (2024)
Adversarial Prompt Tuning for Vision-Language Models
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models
von: Zhou, Qiji, et al.
Veröffentlicht: (2024)
von: Zhou, Qiji, et al.
Veröffentlicht: (2024)
Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
RiskCueBench: Benchmarking Anticipatory Reasoning from Early Risk Cues in Video-Language Models
von: Luo, Sha, et al.
Veröffentlicht: (2026)
von: Luo, Sha, et al.
Veröffentlicht: (2026)
Enhancing Vision-Language Models for Autonomous Driving through Task-Specific Prompting and Spatial Reasoning
von: Wu, Aodi, et al.
Veröffentlicht: (2025)
von: Wu, Aodi, et al.
Veröffentlicht: (2025)
Biomed-DPT: Dual Modality Prompt Tuning for Biomedical Vision-Language Models
von: Peng, Wei, et al.
Veröffentlicht: (2025)
von: Peng, Wei, et al.
Veröffentlicht: (2025)
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
von: Lee, Youngwan, et al.
Veröffentlicht: (2026)
von: Lee, Youngwan, et al.
Veröffentlicht: (2026)
Adversarial Prompt Injection Attack on Multimodal Large Language Models
von: Ding, Meiwen, et al.
Veröffentlicht: (2026)
von: Ding, Meiwen, et al.
Veröffentlicht: (2026)
Document Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documents
von: Chen, Jun, et al.
Veröffentlicht: (2024)
von: Chen, Jun, et al.
Veröffentlicht: (2024)
An Examination of the Compositionality of Large Generative Vision-Language Models
von: Ma, Teli, et al.
Veröffentlicht: (2023)
von: Ma, Teli, et al.
Veröffentlicht: (2023)
Evolving Prompt Adaptation for Vision-Language Models
von: Zhang, Enming, et al.
Veröffentlicht: (2026)
von: Zhang, Enming, et al.
Veröffentlicht: (2026)
Efficient Prompt Tuning of Large Vision-Language Model for Fine-Grained Ship Classification
von: Lan, Long, et al.
Veröffentlicht: (2024)
von: Lan, Long, et al.
Veröffentlicht: (2024)
Gram-Anchored Prompt Learning for Vision-Language Models via Second-Order Statistics
von: Chen, Minglei, et al.
Veröffentlicht: (2026)
von: Chen, Minglei, et al.
Veröffentlicht: (2026)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
von: Che, Liwei, et al.
Veröffentlicht: (2026)
von: Che, Liwei, et al.
Veröffentlicht: (2026)
Target Prompting for Information Extraction with Vision Language Model
von: Medhi, Dipankar
Veröffentlicht: (2024)
von: Medhi, Dipankar
Veröffentlicht: (2024)
Training-Free Unsupervised Prompt for Vision-Language Models
von: Long, Sifan, et al.
Veröffentlicht: (2024)
von: Long, Sifan, et al.
Veröffentlicht: (2024)
Domain-Invariant Prompt Learning for Vision-Language Models
von: Khoee, Arsham Gholamzadeh, et al.
Veröffentlicht: (2026)
von: Khoee, Arsham Gholamzadeh, et al.
Veröffentlicht: (2026)
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
von: Zhao, Bingchen, et al.
Veröffentlicht: (2024)
von: Zhao, Bingchen, et al.
Veröffentlicht: (2024)
MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models
von: Wang, Kangkang, et al.
Veröffentlicht: (2026)
von: Wang, Kangkang, et al.
Veröffentlicht: (2026)
NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language Models
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
LPT: Less-overfitting Prompt Tuning for Vision-Language Model
von: Ding, Chenhao, et al.
Veröffentlicht: (2024)
von: Ding, Chenhao, et al.
Veröffentlicht: (2024)
Tuning Vision-Language Models with Candidate Labels by Prompt Alignment
von: Zhang, Zhifang, et al.
Veröffentlicht: (2024)
von: Zhang, Zhifang, et al.
Veröffentlicht: (2024)
Prompt Tuning with Soft Context Sharing for Vision-Language Models
von: Ding, Kun, et al.
Veröffentlicht: (2022)
von: Ding, Kun, et al.
Veröffentlicht: (2022)
Model Composition for Multimodal Large Language Models
von: Chen, Chi, et al.
Veröffentlicht: (2024)
von: Chen, Chi, et al.
Veröffentlicht: (2024)
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
von: Zhang, Xinsong, et al.
Veröffentlicht: (2025)
von: Zhang, Xinsong, et al.
Veröffentlicht: (2025)
Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models
von: Xing, Wenbin, et al.
Veröffentlicht: (2026)
von: Xing, Wenbin, et al.
Veröffentlicht: (2026)
ReasonEdit: Editing Vision-Language Models using Human Reasoning
von: Qiu, Jiaxing, et al.
Veröffentlicht: (2026)
von: Qiu, Jiaxing, et al.
Veröffentlicht: (2026)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
von: Xiao, Wenyi, et al.
Veröffentlicht: (2025)
von: Xiao, Wenyi, et al.
Veröffentlicht: (2025)
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
PromptEcho: Annotation-Free Reward from Vision-Language Models for Text-to-Image Reinforcement Learning
von: Liu, Jinlong, et al.
Veröffentlicht: (2026)
von: Liu, Jinlong, et al.
Veröffentlicht: (2026)
Delineating Knowledge Boundaries for Honest Large Vision-Language Models
von: Song, Junru, et al.
Veröffentlicht: (2026)
von: Song, Junru, et al.
Veröffentlicht: (2026)
Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models
von: Zhang, Chengsheng, et al.
Veröffentlicht: (2026)
von: Zhang, Chengsheng, et al.
Veröffentlicht: (2026)
Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models
von: Wang, Jiayu, et al.
Veröffentlicht: (2024)
von: Wang, Jiayu, et al.
Veröffentlicht: (2024)
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning
von: Rawal, Ishaan, et al.
Veröffentlicht: (2026)
von: Rawal, Ishaan, et al.
Veröffentlicht: (2026)
A-VL: Adaptive Attention for Large Vision-Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
OLIVE: Object Level In-Context Visual Embeddings
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024) -
Physical Prompt Injection Attacks on Large Vision-Language Models
von: Ling, Chen, et al.
Veröffentlicht: (2026) -
Attention Prompting on Image for Large Vision-Language Models
von: Yu, Runpeng, et al.
Veröffentlicht: (2024) -
Adversarial Prompt Distillation for Vision-Language Models
von: Luo, Lin, et al.
Veröffentlicht: (2024) -
Adversarial Prompt Tuning for Vision-Language Models
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)