Can Modern Vision Models Understand the Difference Between an Object and a Look-alike?
Fuente:
arXiv
Saved in:
| Main Authors: | Cohen, Itay, Fetaya, Ethan, Rosenfeld, Amir |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Questioning the Stability of Visual Question Answering
by: Rosenfeld, Amir, et al.
Published: (2025)
by: Rosenfeld, Amir, et al.
Published: (2025)
Object-Centric Open-Vocabulary Image-Retrieval with Aggregated Features
by: Levi, Hila, et al.
Published: (2023)
by: Levi, Hila, et al.
Published: (2023)
Inverse Problem Sampling in Latent Space Using Sequential Monte Carlo
by: Achituve, Idan, et al.
Published: (2025)
by: Achituve, Idan, et al.
Published: (2025)
From Segments to Concepts: Interpretable Image Classification via Concept-Guided Segmentation
by: Eisenberg, Ran, et al.
Published: (2025)
by: Eisenberg, Ran, et al.
Published: (2025)
Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation
by: Traub, Manuel, et al.
Published: (2025)
by: Traub, Manuel, et al.
Published: (2025)
Towards Understanding Best Practices for Quantization of Vision-Language Models
by: Das, Gautom, et al.
Published: (2026)
by: Das, Gautom, et al.
Published: (2026)
How Can Objects Help Video-Language Understanding?
by: Tang, Zitian, et al.
Published: (2025)
by: Tang, Zitian, et al.
Published: (2025)
ASR: Attention-alike Structural Re-parameterization
by: Zhong, Shanshan, et al.
Published: (2023)
by: Zhong, Shanshan, et al.
Published: (2023)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
Can 3D Vision-Language Models Truly Understand Natural Language?
by: Deng, Weipeng, et al.
Published: (2024)
by: Deng, Weipeng, et al.
Published: (2024)
Can Unified Generation and Understanding Models Maintain Semantic Equivalence Across Different Output Modalities?
by: Jiang, Hongbo, et al.
Published: (2026)
by: Jiang, Hongbo, et al.
Published: (2026)
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
Can Vision Language Models Understand Mimed Actions?
by: Cho, Hyundong, et al.
Published: (2025)
by: Cho, Hyundong, et al.
Published: (2025)
Can Multimodal Large Language Models Truly Understand Small Objects?
by: Han, Fujun, et al.
Published: (2026)
by: Han, Fujun, et al.
Published: (2026)
Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models
by: Luo, Yulin, et al.
Published: (2026)
by: Luo, Yulin, et al.
Published: (2026)
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
by: Yuan, Yuqian, et al.
Published: (2026)
by: Yuan, Yuqian, et al.
Published: (2026)
Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
by: Li, Liyang, et al.
Published: (2026)
by: Li, Liyang, et al.
Published: (2026)
De-Confusing Pseudo-Labels in Source-Free Domain Adaptation
by: Diamant, Idit, et al.
Published: (2024)
by: Diamant, Idit, et al.
Published: (2024)
A Closer Look at Conditional Prompt Tuning for Vision-Language Models
by: Zhang, Ji, et al.
Published: (2025)
by: Zhang, Ji, et al.
Published: (2025)
Can Vision-Language Models Understand Construction Workers? An Exploratory Study
by: Bui, Hieu, et al.
Published: (2026)
by: Bui, Hieu, et al.
Published: (2026)
Show and Tell: Visually Explainable Deep Neural Nets via Spatially-Aware Concept Bottleneck Models
by: Benou, Itay, et al.
Published: (2025)
by: Benou, Itay, et al.
Published: (2025)
BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusion
by: Xiang, Sike, et al.
Published: (2025)
by: Xiang, Sike, et al.
Published: (2025)
VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
by: Zhao, Hongbo, et al.
Published: (2025)
by: Zhao, Hongbo, et al.
Published: (2025)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
by: Min, Cheolhong, et al.
Published: (2026)
by: Min, Cheolhong, et al.
Published: (2026)
A Closer Look at the Few-Shot Adaptation of Large Vision-Language Models
by: Silva-Rodríguez, Julio, et al.
Published: (2023)
by: Silva-Rodríguez, Julio, et al.
Published: (2023)
Vision-Language Models Can't See the Obvious
by: Dahou, Yasser, et al.
Published: (2025)
by: Dahou, Yasser, et al.
Published: (2025)
Look Around and Learn: Self-Training Object Detection by Exploration
by: Scarpellini, Gianluca, et al.
Published: (2023)
by: Scarpellini, Gianluca, et al.
Published: (2023)
Understanding Degradation with Vision Language Model
by: Lan, Guanzhou, et al.
Published: (2026)
by: Lan, Guanzhou, et al.
Published: (2026)
SMART-Vision: Survey of Modern Action Recognition Techniques in Vision
by: AlShami, Ali K., et al.
Published: (2025)
by: AlShami, Ali K., et al.
Published: (2025)
Physical Object Understanding with a Physically Controllable World Model
by: Venkatesh, Rahul, et al.
Published: (2026)
by: Venkatesh, Rahul, et al.
Published: (2026)
Mechanisms of Object Localization in Vision-Language Models
by: Schaumlöffel, Timothy, et al.
Published: (2026)
by: Schaumlöffel, Timothy, et al.
Published: (2026)
SurgCheck: Do Vision-Language Models Really Look at Images in Surgical VQA?
by: Shin, Jongmin, et al.
Published: (2026)
by: Shin, Jongmin, et al.
Published: (2026)
AI-based Wearable Vision Assistance System for the Visually Impaired: Integrating Real-Time Object Recognition and Contextual Understanding Using Large Vision-Language Models
by: Baig, Mirza Samad Ahmed, et al.
Published: (2024)
by: Baig, Mirza Samad Ahmed, et al.
Published: (2024)
A New Hybrid Model of Generative Adversarial Network and You Only Look Once Algorithm for Automatic License-Plate Recognition
by: Shafiezadeh, Behnoud, et al.
Published: (2025)
by: Shafiezadeh, Behnoud, et al.
Published: (2025)
CanViT: Toward Active-Vision Foundation Models
by: Berreby, Yohaï-Eliel, et al.
Published: (2026)
by: Berreby, Yohaï-Eliel, et al.
Published: (2026)
Understanding Dataset Bias in Medical Imaging: A Case Study on Chest X-rays
by: Dack, Ethan, et al.
Published: (2025)
by: Dack, Ethan, et al.
Published: (2025)
Foundations and Models in Modern Computer Vision: Key Building Blocks in Landmark Architectures
by: Bourceanu, Radu-Andrei, et al.
Published: (2025)
by: Bourceanu, Radu-Andrei, et al.
Published: (2025)
Moving by Looking: Towards Vision-Driven Avatar Motion Generation
by: Diomataris, Markos, et al.
Published: (2025)
by: Diomataris, Markos, et al.
Published: (2025)
LookHere: Vision Transformers with Directed Attention Generalize and Extrapolate
by: Fuller, Anthony, et al.
Published: (2024)
by: Fuller, Anthony, et al.
Published: (2024)
Similar Items
-
Questioning the Stability of Visual Question Answering
by: Rosenfeld, Amir, et al.
Published: (2025) -
Object-Centric Open-Vocabulary Image-Retrieval with Aggregated Features
by: Levi, Hila, et al.
Published: (2023) -
Inverse Problem Sampling in Latent Space Using Sequential Monte Carlo
by: Achituve, Idan, et al.
Published: (2025) -
From Segments to Concepts: Interpretable Image Classification via Concept-Guided Segmentation
by: Eisenberg, Ran, et al.
Published: (2025) -
Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation
by: Traub, Manuel, et al.
Published: (2025)