Investigating Mechanisms for In-Context Vision Language Binding
Fuente:
arXiv
Saved in:
| Main Authors: | Saravanan, Darshana, Tapaswi, Makarand, Gandhi, Vineet |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment
by: Saravanan, Darshana, et al.
Published: (2024)
by: Saravanan, Darshana, et al.
Published: (2024)
What You See is What You Ask: Evaluating Audio Descriptions
by: Kala, Divy, et al.
Published: (2025)
by: Kala, Divy, et al.
Published: (2025)
Pseudo-labelling meets Label Smoothing for Noisy Partial Label Learning
by: Saravanan, Darshana, et al.
Published: (2024)
by: Saravanan, Darshana, et al.
Published: (2024)
Steerable Visual Representations
by: Ruthardt, Jona, et al.
Published: (2026)
by: Ruthardt, Jona, et al.
Published: (2026)
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
by: Manikantan, Kawshik, et al.
Published: (2024)
by: Manikantan, Kawshik, et al.
Published: (2024)
MALeR: Improving Compositional Fidelity in Layout-Guided Generation
by: Saxena, Shivank, et al.
Published: (2025)
by: Saxena, Shivank, et al.
Published: (2025)
No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning
by: Gaur, Manu, et al.
Published: (2024)
by: Gaur, Manu, et al.
Published: (2024)
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition
by: Darur, Balaji, et al.
Published: (2026)
by: Darur, Balaji, et al.
Published: (2026)
SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
by: Singh, Darshan, et al.
Published: (2024)
by: Singh, Darshan, et al.
Published: (2024)
Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation
by: Gaur, Manu, et al.
Published: (2024)
by: Gaur, Manu, et al.
Published: (2024)
Major Entity Identification: A Generalizable Alternative to Coreference Resolution
by: Manikantan, Kawshik, et al.
Published: (2024)
by: Manikantan, Kawshik, et al.
Published: (2024)
"Previously on ..." From Recaps to Story Summarization
by: Singh, Aditya Kumar, et al.
Published: (2024)
by: Singh, Aditya Kumar, et al.
Published: (2024)
Seeing Eye to AI: Comparing Human Gaze and Model Attention in Video Memorability
by: Kumar, Prajneya, et al.
Published: (2023)
by: Kumar, Prajneya, et al.
Published: (2023)
Fine-Tuning Without Forgetting: Adaptation of YOLOv8 Preserves COCO Performance
by: Gandhi, Vishal, et al.
Published: (2025)
by: Gandhi, Vishal, et al.
Published: (2025)
Personalized Image Generation from an Author Writing Style
by: Gandhi, Sagar, et al.
Published: (2025)
by: Gandhi, Sagar, et al.
Published: (2025)
Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models
by: Wang, Jiayu, et al.
Published: (2024)
by: Wang, Jiayu, et al.
Published: (2024)
Phantasia: Context-Adaptive Backdoors in Vision Language Models
by: Tran, Nam Duong, et al.
Published: (2026)
by: Tran, Nam Duong, et al.
Published: (2026)
MICap: A Unified Model for Identity-aware Movie Descriptions
by: Raajesh, Haran, et al.
Published: (2024)
by: Raajesh, Haran, et al.
Published: (2024)
Fine-Tuning Vision-Language Models for Visual Navigation Assistance
by: Li, Xiao, et al.
Published: (2025)
by: Li, Xiao, et al.
Published: (2025)
Understanding Counting Mechanisms in Large Language and Vision-Language Models
by: Hasani, Hosein, et al.
Published: (2025)
by: Hasani, Hosein, et al.
Published: (2025)
Prompt Tuning with Soft Context Sharing for Vision-Language Models
by: Ding, Kun, et al.
Published: (2022)
by: Ding, Kun, et al.
Published: (2022)
In-Context Learning Improves Compositional Understanding of Vision-Language Models
by: Nulli, Matteo, et al.
Published: (2024)
by: Nulli, Matteo, et al.
Published: (2024)
Large Vision-Language Models as Emotion Recognizers in Context Awareness
by: Lei, Yuxuan, et al.
Published: (2024)
by: Lei, Yuxuan, et al.
Published: (2024)
BiasICL: In-Context Learning and Demographic Biases of Vision Language Models
by: Xu, Sonnet, et al.
Published: (2025)
by: Xu, Sonnet, et al.
Published: (2025)
Enhancing Agentic Autonomous Scientific Discovery with Vision-Language Model Capabilities
by: Gandhi, Kahaan, et al.
Published: (2025)
by: Gandhi, Kahaan, et al.
Published: (2025)
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
by: Wang, Yizhou, et al.
Published: (2025)
by: Wang, Yizhou, et al.
Published: (2025)
Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models
by: Liu, Yufang, et al.
Published: (2024)
by: Liu, Yufang, et al.
Published: (2024)
GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language Models
by: Abdollahi, Ali, et al.
Published: (2024)
by: Abdollahi, Ali, et al.
Published: (2024)
Object-centric Binding in Contrastive Language-Image Pretraining
by: Assouel, Rim, et al.
Published: (2025)
by: Assouel, Rim, et al.
Published: (2025)
A Generative Approach to High Fidelity 3D Reconstruction from Text Data
by: R, Venkat Kumar, et al.
Published: (2025)
by: R, Venkat Kumar, et al.
Published: (2025)
LongFly: Long-Horizon UAV Vision-and-Language Navigation with Spatiotemporal Context Integration
by: Jiang, Wen, et al.
Published: (2025)
by: Jiang, Wen, et al.
Published: (2025)
Leveraging Chat-Based Large Vision Language Models for Multimodal Out-Of-Context Detection
by: Shalabi, Fatma, et al.
Published: (2024)
by: Shalabi, Fatma, et al.
Published: (2024)
Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction
by: Maeda, Koki, et al.
Published: (2024)
by: Maeda, Koki, et al.
Published: (2024)
Physics Context Builders: A Modular Framework for Physical Reasoning in Vision-Language Models
by: Balazadeh, Vahid, et al.
Published: (2024)
by: Balazadeh, Vahid, et al.
Published: (2024)
I Walk the Line: Examining the Role of Gestalt Continuity in Object Binding for Vision Transformers
by: Tartaglini, Alexa R., et al.
Published: (2026)
by: Tartaglini, Alexa R., et al.
Published: (2026)
VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
by: Zhao, Hongbo, et al.
Published: (2025)
by: Zhao, Hongbo, et al.
Published: (2025)
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
by: Zhu, Bin, et al.
Published: (2023)
by: Zhu, Bin, et al.
Published: (2023)
Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models
by: Panchal, Utsav, et al.
Published: (2025)
by: Panchal, Utsav, et al.
Published: (2025)
Using Vision Language Foundation Models to Generate Plant Simulation Configurations via In-Context Learning
by: Yun, Heesup, et al.
Published: (2026)
by: Yun, Heesup, et al.
Published: (2026)
CARPE: Context-Aware Image Representation Prioritization via Ensemble for Large Vision-Language Models
by: Lee, Donghee, et al.
Published: (2026)
by: Lee, Donghee, et al.
Published: (2026)
Similar Items
-
VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment
by: Saravanan, Darshana, et al.
Published: (2024) -
What You See is What You Ask: Evaluating Audio Descriptions
by: Kala, Divy, et al.
Published: (2025) -
Pseudo-labelling meets Label Smoothing for Noisy Partial Label Learning
by: Saravanan, Darshana, et al.
Published: (2024) -
Steerable Visual Representations
by: Ruthardt, Jona, et al.
Published: (2026) -
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
by: Manikantan, Kawshik, et al.
Published: (2024)