Hidden in plain sight: VLMs overlook their visual representations
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Stephanie, Bonnen, Tyler, Guillory, Devin, Darrell, Trevor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
Hyperbolic Active Learning for Semantic Segmentation under Domain Shift
by: Franco, Luca, et al.
Published: (2023)
by: Franco, Luca, et al.
Published: (2023)
UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity
by: Yu, Junwei, et al.
Published: (2025)
by: Yu, Junwei, et al.
Published: (2025)
Reconstruction Alignment Improves Unified Multimodal Models
by: Xie, Ji, et al.
Published: (2025)
by: Xie, Ji, et al.
Published: (2025)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
by: Mitra, Chancharik, et al.
Published: (2023)
by: Mitra, Chancharik, et al.
Published: (2023)
REOrdering Patches Improves Vision Models
by: Kutscher, Declan, et al.
Published: (2025)
by: Kutscher, Declan, et al.
Published: (2025)
Visual Lexicon: Rich Image Features in Language Space
by: Wang, XuDong, et al.
Published: (2024)
by: Wang, XuDong, et al.
Published: (2024)
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024)
by: Wang, Xudong, et al.
Published: (2024)
Lifting Embodied World Models for Planning and Control
by: Wang, Alex N., et al.
Published: (2026)
by: Wang, Alex N., et al.
Published: (2026)
Galaxy Walker: Geometry-aware VLMs For Galaxy-scale Understanding
by: Chen, Tianyu, et al.
Published: (2025)
by: Chen, Tianyu, et al.
Published: (2025)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
by: Shang, Chuyi, et al.
Published: (2024)
by: Shang, Chuyi, et al.
Published: (2024)
Navigation World Models
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
Rethinking VLMs and LLMs for Image Classification
by: Cooper, Avi, et al.
Published: (2024)
by: Cooper, Avi, et al.
Published: (2024)
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
by: Li, Shuo, et al.
Published: (2024)
by: Li, Shuo, et al.
Published: (2024)
Self-supervised video pretraining yields robust and more human-aligned visual representations
by: Parthasarathy, Nikhil, et al.
Published: (2022)
by: Parthasarathy, Nikhil, et al.
Published: (2022)
CARES: Context-Aware Resolution Selector for VLMs
by: Kimhi, Moshe, et al.
Published: (2025)
by: Kimhi, Moshe, et al.
Published: (2025)
DASH: Detection and Assessment of Systematic Hallucinations of VLMs
by: Augustin, Maximilian, et al.
Published: (2025)
by: Augustin, Maximilian, et al.
Published: (2025)
Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
by: Moon, Jihyun, et al.
Published: (2025)
by: Moon, Jihyun, et al.
Published: (2025)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
by: Huang, Brandon, et al.
Published: (2024)
by: Huang, Brandon, et al.
Published: (2024)
Video Action Differencing
by: Burgess, James, et al.
Published: (2025)
by: Burgess, James, et al.
Published: (2025)
Shape-Guided Diffusion with Inside-Outside Attention
by: Park, Dong Huk, et al.
Published: (2022)
by: Park, Dong Huk, et al.
Published: (2022)
Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight
by: Ding, Xi, et al.
Published: (2024)
by: Ding, Xi, et al.
Published: (2024)
DEAL: Disentangle and Localize Concept-level Explanations for VLMs
by: Li, Tang, et al.
Published: (2024)
by: Li, Tang, et al.
Published: (2024)
PAIR-Diffusion: A Comprehensive Multimodal Object-Level Image Editor
by: Goel, Vidit, et al.
Published: (2023)
by: Goel, Vidit, et al.
Published: (2023)
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
by: Faraz, Ali, et al.
Published: (2025)
by: Faraz, Ali, et al.
Published: (2025)
Do We Need Large VLMs for Spotting Soccer Actions?
by: Chakraborty, Ritabrata, et al.
Published: (2025)
by: Chakraborty, Ritabrata, et al.
Published: (2025)
ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs
by: Burgess, James, et al.
Published: (2026)
by: Burgess, James, et al.
Published: (2026)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
Can VLMs Reason Robustly? A Neuro-Symbolic Investigation
by: Chen, Weixin, et al.
Published: (2026)
by: Chen, Weixin, et al.
Published: (2026)
Stochastic positional embeddings improve masked image modeling
by: Bar, Amir, et al.
Published: (2023)
by: Bar, Amir, et al.
Published: (2023)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
by: Lee, Heekyung, et al.
Published: (2025)
by: Lee, Heekyung, et al.
Published: (2025)
Analyzing The Language of Visual Tokens
by: Chan, David M., et al.
Published: (2024)
by: Chan, David M., et al.
Published: (2024)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
by: Izadi, Amirmohammad, et al.
Published: (2025)
by: Izadi, Amirmohammad, et al.
Published: (2025)
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
by: Li, Kevin Y., et al.
Published: (2024)
by: Li, Kevin Y., et al.
Published: (2024)
COVR:Collaborative Optimization of VLMs and RL Agent for Visual-Based Control
by: Xia, Canming, et al.
Published: (2026)
by: Xia, Canming, et al.
Published: (2026)
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
by: Kumar, Sunil, et al.
Published: (2025)
by: Kumar, Sunil, et al.
Published: (2025)
Right this way: Can VLMs Guide Us to See More to Answer Questions?
by: Liu, Li, et al.
Published: (2024)
by: Liu, Li, et al.
Published: (2024)
ALOHa: A New Measure for Hallucination in Captioning Models
by: Petryk, Suzanne, et al.
Published: (2024)
by: Petryk, Suzanne, et al.
Published: (2024)
Whole-Body Conditioned Egocentric Video Prediction
by: Bai, Yutong, et al.
Published: (2025)
by: Bai, Yutong, et al.
Published: (2025)
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
by: Kim, Minkyu, et al.
Published: (2026)
by: Kim, Minkyu, et al.
Published: (2026)
Similar Items
-
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025) -
Hyperbolic Active Learning for Semantic Segmentation under Domain Shift
by: Franco, Luca, et al.
Published: (2023) -
UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity
by: Yu, Junwei, et al.
Published: (2025) -
Reconstruction Alignment Improves Unified Multimodal Models
by: Xie, Ji, et al.
Published: (2025) -
Compositional Chain-of-Thought Prompting for Large Multimodal Models
by: Mitra, Chancharik, et al.
Published: (2023)