Lost in the Vibrations: Vision Language Models Fail the Dynamic Gauges Test
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fu, Tairan, Santos-Martín, Francisco Javier, Conde, Javier, Reviriego, Pedro, Merino-Gómez, Elena |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Recursive InPainting (RIP): how much information is lost under recursive inferences?
von: Conde, Javier, et al.
Veröffentlicht: (2024)
von: Conde, Javier, et al.
Veröffentlicht: (2024)
Have Multimodal Large Language Models (MLLMs) Really Learned to Tell the Time on Analog Clocks?
von: Fu, Tairan, et al.
Veröffentlicht: (2025)
von: Fu, Tairan, et al.
Veröffentlicht: (2025)
Large Language Models and Book Summarization: Reading or Remembering, Which Is Better?
von: Fu, Tairan, et al.
Veröffentlicht: (2026)
von: Fu, Tairan, et al.
Veröffentlicht: (2026)
Evaluating Large Language Models with Tests of Spanish as a Foreign Language: Pass or Fail?
von: Mayor-Rocher, Marina, et al.
Veröffentlicht: (2024)
von: Mayor-Rocher, Marina, et al.
Veröffentlicht: (2024)
A Lost Opportunity for Vision-Language Models: A Comparative Study of Online Test-Time Adaptation for Vision-Language Models
von: Döbler, Mario, et al.
Veröffentlicht: (2024)
von: Döbler, Mario, et al.
Veröffentlicht: (2024)
Lost in Embeddings: Information Loss in Vision-Language Models
von: Li, Wenyan, et al.
Veröffentlicht: (2025)
von: Li, Wenyan, et al.
Veröffentlicht: (2025)
Large Vision-Language Models Get Lost in Attention
von: Xi, Gongli, et al.
Veröffentlicht: (2026)
von: Xi, Gongli, et al.
Veröffentlicht: (2026)
Texture or Semantics? Vision-Language Models Get Lost in Font Recognition
von: Li, Zhecheng, et al.
Veröffentlicht: (2025)
von: Li, Zhecheng, et al.
Veröffentlicht: (2025)
Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?
von: Qu, Tianyuan, et al.
Veröffentlicht: (2025)
von: Qu, Tianyuan, et al.
Veröffentlicht: (2025)
Color Names in Vision-Language Models
von: Gomez-Villa, Alexandra, et al.
Veröffentlicht: (2025)
von: Gomez-Villa, Alexandra, et al.
Veröffentlicht: (2025)
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
von: Awad, Samer, et al.
Veröffentlicht: (2026)
von: Awad, Samer, et al.
Veröffentlicht: (2026)
Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization
von: Bharadwaj, Siddhant, et al.
Veröffentlicht: (2026)
von: Bharadwaj, Siddhant, et al.
Veröffentlicht: (2026)
Your other Left! Vision-Language Models Fail to Identify Relative Positions in Medical Images
von: Wolf, Daniel, et al.
Veröffentlicht: (2025)
von: Wolf, Daniel, et al.
Veröffentlicht: (2025)
Cross-Attentive Multiview Fusion of Vision-Language Embeddings
von: Martins, Tomas Berriel, et al.
Veröffentlicht: (2026)
von: Martins, Tomas Berriel, et al.
Veröffentlicht: (2026)
Why Do Large Language Models (LLMs) Struggle to Count Letters?
von: Fu, Tairan, et al.
Veröffentlicht: (2024)
von: Fu, Tairan, et al.
Veröffentlicht: (2024)
When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
von: Yan, Yuping, et al.
Veröffentlicht: (2025)
von: Yan, Yuping, et al.
Veröffentlicht: (2025)
Test-Time Consistency in Vision Language Models
von: Chou, Shih-Han, et al.
Veröffentlicht: (2025)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2025)
Artificial Intelligence and Misinformation in Art: Can Vision Language Models Judge the Hand or the Machine Behind the Canvas?
von: Fu, Tarian, et al.
Veröffentlicht: (2025)
von: Fu, Tarian, et al.
Veröffentlicht: (2025)
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation
von: Deng, Ken, et al.
Veröffentlicht: (2026)
von: Deng, Ken, et al.
Veröffentlicht: (2026)
ETTA: Efficient Test-Time Adaptation for Vision-Language Models through Dynamic Embedding Updates
von: Dastmalchi, Hamidreza, et al.
Veröffentlicht: (2025)
von: Dastmalchi, Hamidreza, et al.
Veröffentlicht: (2025)
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
von: Fu, Tairan, et al.
Veröffentlicht: (2025)
von: Fu, Tairan, et al.
Veröffentlicht: (2025)
Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models
von: Monon, Mashrafi, et al.
Veröffentlicht: (2026)
von: Monon, Mashrafi, et al.
Veröffentlicht: (2026)
Realistic Test-Time Adaptation of Vision-Language Models
von: Zanella, Maxime, et al.
Veröffentlicht: (2025)
von: Zanella, Maxime, et al.
Veröffentlicht: (2025)
Bayesian Test-Time Adaptation for Vision-Language Models
von: Zhou, Lihua, et al.
Veröffentlicht: (2025)
von: Zhou, Lihua, et al.
Veröffentlicht: (2025)
Efficient Test-Time Adaptation of Vision-Language Models
von: Karmanov, Adilbek, et al.
Veröffentlicht: (2024)
von: Karmanov, Adilbek, et al.
Veröffentlicht: (2024)
Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models
von: Xia, Hou, et al.
Veröffentlicht: (2025)
von: Xia, Hou, et al.
Veröffentlicht: (2025)
Evaluation of Vision Transformers for Multimodal Image Classification: A Case Study on Brain, Lung, and Kidney Tumors
von: Martín, Óscar A., et al.
Veröffentlicht: (2025)
von: Martín, Óscar A., et al.
Veröffentlicht: (2025)
Test-Time Hinting for Black-Box Vision-Language Models
von: Hou, Kaihua, et al.
Veröffentlicht: (2026)
von: Hou, Kaihua, et al.
Veröffentlicht: (2026)
Ultra-Light Test-Time Adaptation for Vision--Language Models
von: Kim, Byunghyun
Veröffentlicht: (2025)
von: Kim, Byunghyun
Veröffentlicht: (2025)
Prototype-Based Test-Time Adaptation of Vision-Language Models
von: Huang, Zhaohong, et al.
Veröffentlicht: (2026)
von: Huang, Zhaohong, et al.
Veröffentlicht: (2026)
Efficient Test-Time Prompt Tuning for Vision-Language Models
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
Online Gaussian Test-Time Adaptation of Vision-Language Models
von: Fuchs, Clément, et al.
Veröffentlicht: (2025)
von: Fuchs, Clément, et al.
Veröffentlicht: (2025)
Negation-Aware Test-Time Adaptation for Vision-Language Models
von: Han, Haochen, et al.
Veröffentlicht: (2025)
von: Han, Haochen, et al.
Veröffentlicht: (2025)
TTRV: Test-Time Reinforcement Learning for Vision Language Models
von: Singh, Akshit, et al.
Veröffentlicht: (2025)
von: Singh, Akshit, et al.
Veröffentlicht: (2025)
Flatness Guided Test-Time Adaptation for Vision-Language Models
von: Li, Aodi, et al.
Veröffentlicht: (2025)
von: Li, Aodi, et al.
Veröffentlicht: (2025)
Test-time Alignment-Enhanced Adapter for Vision-Language Models
von: Tong, Baoshun, et al.
Veröffentlicht: (2024)
von: Tong, Baoshun, et al.
Veröffentlicht: (2024)
Dynamic Rank Adaptation for Vision-Language Models
von: Wang, Jiahui, et al.
Veröffentlicht: (2025)
von: Wang, Jiahui, et al.
Veröffentlicht: (2025)
Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language Resamplers
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2024)
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2024)
Proxy Robustness in Vision Language Models is Effortlessly Transferable
von: Fu, Xiaowei, et al.
Veröffentlicht: (2026)
von: Fu, Xiaowei, et al.
Veröffentlicht: (2026)
Self-Correction Inside the Model: Leveraging Layer Attention to Mitigate Hallucinations in Large Vision Language Models
von: Fu, April
Veröffentlicht: (2026)
von: Fu, April
Veröffentlicht: (2026)
Ähnliche Einträge
-
Recursive InPainting (RIP): how much information is lost under recursive inferences?
von: Conde, Javier, et al.
Veröffentlicht: (2024) -
Have Multimodal Large Language Models (MLLMs) Really Learned to Tell the Time on Analog Clocks?
von: Fu, Tairan, et al.
Veröffentlicht: (2025) -
Large Language Models and Book Summarization: Reading or Remembering, Which Is Better?
von: Fu, Tairan, et al.
Veröffentlicht: (2026) -
Evaluating Large Language Models with Tests of Spanish as a Foreign Language: Pass or Fail?
von: Mayor-Rocher, Marina, et al.
Veröffentlicht: (2024) -
A Lost Opportunity for Vision-Language Models: A Comparative Study of Online Test-Time Adaptation for Vision-Language Models
von: Döbler, Mario, et al.
Veröffentlicht: (2024)