Mechanisms of Object Localization in Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Schaumlöffel, Timothy, Vilas, Martina G., Roig, Gemma |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Contextual inference from single objects in Vision-Language models
von: Vilas, Martina G., et al.
Veröffentlicht: (2026)
von: Vilas, Martina G., et al.
Veröffentlicht: (2026)
Temporal Slowness in Central Vision Drives Semantic Object Learning
von: Schaumlöffel, Timothy, et al.
Veröffentlicht: (2026)
von: Schaumlöffel, Timothy, et al.
Veröffentlicht: (2026)
Human Gaze Boosts Object-Centered Representation Learning
von: Schaumlöffel, Timothy, et al.
Veröffentlicht: (2025)
von: Schaumlöffel, Timothy, et al.
Veröffentlicht: (2025)
Learning Object Semantic Similarity with Self-Supervision
von: Aubret, Arthur, et al.
Veröffentlicht: (2024)
von: Aubret, Arthur, et al.
Veröffentlicht: (2024)
Caregiver Talk Shapes Toddler Vision: A Computational Study of Dyadic Play
von: Schaumlöffel, Timothy, et al.
Veröffentlicht: (2023)
von: Schaumlöffel, Timothy, et al.
Veröffentlicht: (2023)
Evaluation of Randomization through Style Transfer for Enhanced Domain Generalization
von: Eisenhardt, Dustin, et al.
Veröffentlicht: (2026)
von: Eisenhardt, Dustin, et al.
Veröffentlicht: (2026)
FovEx: Human-Inspired Explanations for Vision Transformers and Convolutional Neural Networks
von: Panda, Mahadev Prasad, et al.
Veröffentlicht: (2024)
von: Panda, Mahadev Prasad, et al.
Veröffentlicht: (2024)
MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
von: Galella, Santiago, et al.
Veröffentlicht: (2026)
von: Galella, Santiago, et al.
Veröffentlicht: (2026)
Net2Brain: A Toolbox to compare artificial vision models with human brain responses
von: Bersch, Domenic, et al.
Veröffentlicht: (2022)
von: Bersch, Domenic, et al.
Veröffentlicht: (2022)
Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization
von: Hannan, Darryl, et al.
Veröffentlicht: (2025)
von: Hannan, Darryl, et al.
Veröffentlicht: (2025)
Locality Alignment Improves Vision-Language Models
von: Covert, Ian, et al.
Veröffentlicht: (2024)
von: Covert, Ian, et al.
Veröffentlicht: (2024)
Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation
von: Traub, Manuel, et al.
Veröffentlicht: (2025)
von: Traub, Manuel, et al.
Veröffentlicht: (2025)
Dual-Pathway Circuits of Object Hallucination in Vision-Language Models
von: Liu, Jiaxin, et al.
Veröffentlicht: (2026)
von: Liu, Jiaxin, et al.
Veröffentlicht: (2026)
Object Hallucination-Free Reinforcement Unlearning for Vision-Language Models
von: Jia, Kaidi, et al.
Veröffentlicht: (2026)
von: Jia, Kaidi, et al.
Veröffentlicht: (2026)
Revisiting Few-Shot Object Detection with Vision-Language Models
von: Madan, Anish, et al.
Veröffentlicht: (2023)
von: Madan, Anish, et al.
Veröffentlicht: (2023)
Temporal-Spatial Object Relations Modeling for Vision-and-Language Navigation
von: Huang, Bowen, et al.
Veröffentlicht: (2024)
von: Huang, Bowen, et al.
Veröffentlicht: (2024)
Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models
von: Kang, Raphi, et al.
Veröffentlicht: (2026)
von: Kang, Raphi, et al.
Veröffentlicht: (2026)
Training-Free Semantic Multi-Object Tracking with Vision-Language Models
von: Bonat, Laurence, et al.
Veröffentlicht: (2026)
von: Bonat, Laurence, et al.
Veröffentlicht: (2026)
Open-Vocabulary Camouflaged Object Segmentation with Cascaded Vision Language Models
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
Anomaly Object Segmentation with Vision-Language Models for Steel Scrap Recycling
von: Tanaka, Daichi, et al.
Veröffentlicht: (2025)
von: Tanaka, Daichi, et al.
Veröffentlicht: (2025)
A Review of 3D Object Detection with Vision-Language Models
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
CLIP-Clique: Graph-based Correspondence Matching Augmented by Vision Language Models for Object-based Global Localization
von: Matsuzaki, Shigemichi, et al.
Veröffentlicht: (2024)
von: Matsuzaki, Shigemichi, et al.
Veröffentlicht: (2024)
MarvelOVD: Marrying Object Recognition and Vision-Language Models for Robust Open-Vocabulary Object Detection
von: Wang, Kuo, et al.
Veröffentlicht: (2024)
von: Wang, Kuo, et al.
Veröffentlicht: (2024)
Masked Diffusion Vision-Language Models for Temporal Action Localization
von: Wang, Fengshun, et al.
Veröffentlicht: (2026)
von: Wang, Fengshun, et al.
Veröffentlicht: (2026)
Learning Invariant Causal Mechanism from Vision-Language Models
von: Song, Zeen, et al.
Veröffentlicht: (2024)
von: Song, Zeen, et al.
Veröffentlicht: (2024)
Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models
von: Imran, Muhammad, et al.
Veröffentlicht: (2025)
von: Imran, Muhammad, et al.
Veröffentlicht: (2025)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
von: An, Wenbin, et al.
Veröffentlicht: (2024)
von: An, Wenbin, et al.
Veröffentlicht: (2024)
Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models
von: Lovenia, Holy, et al.
Veröffentlicht: (2023)
von: Lovenia, Holy, et al.
Veröffentlicht: (2023)
Object-Centric Vision Token Pruning for Vision Language Models
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
von: Jahagirdar, Soumya, et al.
Veröffentlicht: (2026)
von: Jahagirdar, Soumya, et al.
Veröffentlicht: (2026)
Object Referring-Guided Scanpath Prediction with Perception-Enhanced Vision-Language Models
von: Quan, Rong, et al.
Veröffentlicht: (2026)
von: Quan, Rong, et al.
Veröffentlicht: (2026)
UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language Models
von: Pi, Xinyu, et al.
Veröffentlicht: (2024)
von: Pi, Xinyu, et al.
Veröffentlicht: (2024)
BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
von: Bhat, Vineet, et al.
Veröffentlicht: (2025)
von: Bhat, Vineet, et al.
Veröffentlicht: (2025)
The Spatial Blindspot of Vision-Language Models
von: Alam, Nahid, et al.
Veröffentlicht: (2026)
von: Alam, Nahid, et al.
Veröffentlicht: (2026)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
von: Chen, Shiming, et al.
Veröffentlicht: (2025)
von: Chen, Shiming, et al.
Veröffentlicht: (2025)
HalLoc: Token-level Localization of Hallucinations for Vision Language Models
von: Park, Eunkyu, et al.
Veröffentlicht: (2025)
von: Park, Eunkyu, et al.
Veröffentlicht: (2025)
Articulate-Anything: Automatic Modeling of Articulated Objects via a Vision-Language Foundation Model
von: Le, Long, et al.
Veröffentlicht: (2024)
von: Le, Long, et al.
Veröffentlicht: (2024)
VORD: Visual Ordinal Calibration for Mitigating Object Hallucinations in Large Vision-Language Models
von: Neo, Dexter, et al.
Veröffentlicht: (2024)
von: Neo, Dexter, et al.
Veröffentlicht: (2024)
TaskCLIP: Extend Large Vision-Language Model for Task Oriented Object Detection
von: Chen, Hanning, et al.
Veröffentlicht: (2024)
von: Chen, Hanning, et al.
Veröffentlicht: (2024)
WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation
von: Nie, Dujun, et al.
Veröffentlicht: (2025)
von: Nie, Dujun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Contextual inference from single objects in Vision-Language models
von: Vilas, Martina G., et al.
Veröffentlicht: (2026) -
Temporal Slowness in Central Vision Drives Semantic Object Learning
von: Schaumlöffel, Timothy, et al.
Veröffentlicht: (2026) -
Human Gaze Boosts Object-Centered Representation Learning
von: Schaumlöffel, Timothy, et al.
Veröffentlicht: (2025) -
Learning Object Semantic Similarity with Self-Supervision
von: Aubret, Arthur, et al.
Veröffentlicht: (2024) -
Caregiver Talk Shapes Toddler Vision: A Computational Study of Dyadic Play
von: Schaumlöffel, Timothy, et al.
Veröffentlicht: (2023)