Contextual inference from single objects in Vision-Language models
Fuente:
arXiv
Saved in:
| Main Authors: | Vilas, Martina G., Schaumlöffel, Timothy, Roig, Gemma |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mechanisms of Object Localization in Vision-Language Models
by: Schaumlöffel, Timothy, et al.
Published: (2026)
by: Schaumlöffel, Timothy, et al.
Published: (2026)
Caregiver Talk Shapes Toddler Vision: A Computational Study of Dyadic Play
by: Schaumlöffel, Timothy, et al.
Published: (2023)
by: Schaumlöffel, Timothy, et al.
Published: (2023)
Evaluation of Randomization through Style Transfer for Enhanced Domain Generalization
by: Eisenhardt, Dustin, et al.
Published: (2026)
by: Eisenhardt, Dustin, et al.
Published: (2026)
Temporal Slowness in Central Vision Drives Semantic Object Learning
by: Schaumlöffel, Timothy, et al.
Published: (2026)
by: Schaumlöffel, Timothy, et al.
Published: (2026)
Net2Brain: A Toolbox to compare artificial vision models with human brain responses
by: Bersch, Domenic, et al.
Published: (2022)
by: Bersch, Domenic, et al.
Published: (2022)
Human Gaze Boosts Object-Centered Representation Learning
by: Schaumlöffel, Timothy, et al.
Published: (2025)
by: Schaumlöffel, Timothy, et al.
Published: (2025)
Learning Object Semantic Similarity with Self-Supervision
by: Aubret, Arthur, et al.
Published: (2024)
by: Aubret, Arthur, et al.
Published: (2024)
CoMa: Contextual Massing Generation with Vision-Language Models
by: Maslov, Evgenii, et al.
Published: (2026)
by: Maslov, Evgenii, et al.
Published: (2026)
ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models
by: Yi, Jingwei, et al.
Published: (2025)
by: Yi, Jingwei, et al.
Published: (2025)
Prompting Large Vision-Language Models for Compositional Reasoning
by: Ossowski, Timothy, et al.
Published: (2024)
by: Ossowski, Timothy, et al.
Published: (2024)
Vision language models have difficulty recognizing virtual objects
by: Tran, Tyler, et al.
Published: (2025)
by: Tran, Tyler, et al.
Published: (2025)
FovEx: Human-Inspired Explanations for Vision Transformers and Convolutional Neural Networks
by: Panda, Mahadev Prasad, et al.
Published: (2024)
by: Panda, Mahadev Prasad, et al.
Published: (2024)
Contextual Object Detection with Multimodal Large Language Models
by: Zang, Yuhang, et al.
Published: (2023)
by: Zang, Yuhang, et al.
Published: (2023)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
by: Tang, Zicong, et al.
Published: (2025)
by: Tang, Zicong, et al.
Published: (2025)
On Explaining Knowledge Distillation: Measuring and Visualising the Knowledge Transfer Process
by: Adhane, Gereziher, et al.
Published: (2024)
by: Adhane, Gereziher, et al.
Published: (2024)
Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
by: Chatzoudis, Gerasimos, et al.
Published: (2026)
by: Chatzoudis, Gerasimos, et al.
Published: (2026)
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
by: Jung, Woojun, et al.
Published: (2025)
by: Jung, Woojun, et al.
Published: (2025)
SCOPE: Sign Language Contextual Processing with Embedding from LLMs
by: Liu, Yuqi, et al.
Published: (2024)
by: Liu, Yuqi, et al.
Published: (2024)
Quantum Inverse Contextual Vision Transformers (Q-ICVT): A New Frontier in 3D Object Detection for AVs
by: Dharavath, Sanjay Bhargav, et al.
Published: (2024)
by: Dharavath, Sanjay Bhargav, et al.
Published: (2024)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
by: Zhang, Jianke, et al.
Published: (2026)
by: Zhang, Jianke, et al.
Published: (2026)
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
by: Saini, Harshvardhan, et al.
Published: (2026)
by: Saini, Harshvardhan, et al.
Published: (2026)
Continual Vision-and-Language Navigation
by: Jeong, Seongjun, et al.
Published: (2024)
by: Jeong, Seongjun, et al.
Published: (2024)
Object-Centric Vision Token Pruning for Vision Language Models
by: Li, Guangyuan, et al.
Published: (2025)
by: Li, Guangyuan, et al.
Published: (2025)
Intersectional Fairness in Vision-Language Models for Medical Image Disease Classification
by: Zhang, Yupeng, et al.
Published: (2025)
by: Zhang, Yupeng, et al.
Published: (2025)
BabyFlow: 3D modeling of realistic and expressive infant faces
by: Alomar, Antonia, et al.
Published: (2025)
by: Alomar, Antonia, et al.
Published: (2025)
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
by: Zhao, Bingchen, et al.
Published: (2024)
by: Zhao, Bingchen, et al.
Published: (2024)
Language-Guided Invariance Probing of Vision-Language Models
by: Lee, Jae Joong
Published: (2025)
by: Lee, Jae Joong
Published: (2025)
AYDIV: Adaptable Yielding 3D Object Detection via Integrated Contextual Vision Transformer
by: Dam, Tanmoy, et al.
Published: (2024)
by: Dam, Tanmoy, et al.
Published: (2024)
Visual moral inference and communication
by: Zhu, Warren, et al.
Published: (2025)
by: Zhu, Warren, et al.
Published: (2025)
CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
by: Stilz, Florian, et al.
Published: (2026)
by: Stilz, Florian, et al.
Published: (2026)
What's in the Image? A Deep-Dive into the Vision of Vision Language Models
by: Kaduri, Omri, et al.
Published: (2024)
by: Kaduri, Omri, et al.
Published: (2024)
LRR-Bench: Left, Right or Rotate? Vision-Language models Still Struggle With Spatial Understanding Tasks
by: Kong, Fei, et al.
Published: (2025)
by: Kong, Fei, et al.
Published: (2025)
FFF: Fixing Flawed Foundations in contrastive pre-training results in very strong Vision-Language models
by: Bulat, Adrian, et al.
Published: (2024)
by: Bulat, Adrian, et al.
Published: (2024)
IoT Botnet Detection: Application of Vision Transformer to Classification of Network Flow Traffic
by: Wasswa, Hassan, et al.
Published: (2025)
by: Wasswa, Hassan, et al.
Published: (2025)
Egocentric Bias in Vision-Language Models
by: Wang, Maijunxian, et al.
Published: (2026)
by: Wang, Maijunxian, et al.
Published: (2026)
Enhance Vision-Language Alignment with Noise
by: Huang, Sida, et al.
Published: (2024)
by: Huang, Sida, et al.
Published: (2024)
Safety Alignment for Vision Language Models
by: Liu, Zhendong, et al.
Published: (2024)
by: Liu, Zhendong, et al.
Published: (2024)
Seeing the Abstract: Translating the Abstract Language for Vision Language Models
by: Talon, Davide, et al.
Published: (2025)
by: Talon, Davide, et al.
Published: (2025)
Understanding Counting Mechanisms in Large Language and Vision-Language Models
by: Hasani, Hosein, et al.
Published: (2025)
by: Hasani, Hosein, et al.
Published: (2025)
Are Large Vision-Language Models Ready to Guide Blind and Low-Vision Individuals?
by: Kim, Eunki, et al.
Published: (2025)
by: Kim, Eunki, et al.
Published: (2025)
Similar Items
-
Mechanisms of Object Localization in Vision-Language Models
by: Schaumlöffel, Timothy, et al.
Published: (2026) -
Caregiver Talk Shapes Toddler Vision: A Computational Study of Dyadic Play
by: Schaumlöffel, Timothy, et al.
Published: (2023) -
Evaluation of Randomization through Style Transfer for Enhanced Domain Generalization
by: Eisenhardt, Dustin, et al.
Published: (2026) -
Temporal Slowness in Central Vision Drives Semantic Object Learning
by: Schaumlöffel, Timothy, et al.
Published: (2026) -
Net2Brain: A Toolbox to compare artificial vision models with human brain responses
by: Bersch, Domenic, et al.
Published: (2022)