Decoupling the components of geometric understanding in Vision Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kosoy, Eliza, Dahmani, Annya, Lampinen, Andrew K., Comsa, Iulia M., Jeong, Soojin, Dasgupta, Ishita, Allen, Kelsey |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The in-context inductive biases of vision-language models differ across modalities
by: Allen, Kelsey, et al.
Published: (2025)
by: Allen, Kelsey, et al.
Published: (2025)
Libra: Building Decoupled Vision System on Large Language Models
by: Xu, Yifan, et al.
Published: (2024)
by: Xu, Yifan, et al.
Published: (2024)
Hierarchical Dual-Subspace Decoupling for Continual Learning in Vision-Language Models
by: Qin, Mengxin, et al.
Published: (2026)
by: Qin, Mengxin, et al.
Published: (2026)
Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models
by: Ma, Kexin, et al.
Published: (2026)
by: Ma, Kexin, et al.
Published: (2026)
Context-Based Visual-Language Place Recognition
by: Woo, Soojin, et al.
Published: (2024)
by: Woo, Soojin, et al.
Published: (2024)
3DSPA: A 3D Semantic Point Autoencoder for Evaluating Video Realism
by: Chandna, Bhavik, et al.
Published: (2026)
by: Chandna, Bhavik, et al.
Published: (2026)
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
by: Lee, Youngwan, et al.
Published: (2025)
by: Lee, Youngwan, et al.
Published: (2025)
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
by: Lee, Youngwan, et al.
Published: (2026)
by: Lee, Youngwan, et al.
Published: (2026)
Unified Generation and Self-Verification for Vision-Language Models via Advantage Decoupled Preference Optimization
by: Qiu, Xinyu, et al.
Published: (2026)
by: Qiu, Xinyu, et al.
Published: (2026)
Learned feature representations are biased by complexity, learning order, position, and more
by: Lampinen, Andrew Kyle, et al.
Published: (2024)
by: Lampinen, Andrew Kyle, et al.
Published: (2024)
DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
Zero-Shot Distracted Driver Detection via Vision Language Models with Double Decoupling
by: Miyata, Takamichi, et al.
Published: (2026)
by: Miyata, Takamichi, et al.
Published: (2026)
Vision-aligned Latent Reasoning for Multi-modal Large Language Model
by: Jeon, Byungwoo, et al.
Published: (2026)
by: Jeon, Byungwoo, et al.
Published: (2026)
HanDiffuser: Text-to-Image Generation With Realistic Hand Appearances
by: Narasimhaswamy, Supreeth, et al.
Published: (2024)
by: Narasimhaswamy, Supreeth, et al.
Published: (2024)
Decoupling Augmentation Bias in Prompt Learning for Vision-Language Models
by: Kim, Gahyeon, et al.
Published: (2025)
by: Kim, Gahyeon, et al.
Published: (2025)
ArtGPT-4: Towards Artistic-understanding Large Vision-Language Models with Enhanced Adapter
by: Yuan, Zhengqing, et al.
Published: (2023)
by: Yuan, Zhengqing, et al.
Published: (2023)
SIA: Enhancing Safety via Intent Awareness for Vision-Language Models
by: Na, Youngjin, et al.
Published: (2025)
by: Na, Youngjin, et al.
Published: (2025)
SWAG: Splatting in the Wild images with Appearance-conditioned Gaussians
by: Dahmani, Hiba, et al.
Published: (2024)
by: Dahmani, Hiba, et al.
Published: (2024)
Understanding Visual Feature Reliance through the Lens of Complexity
by: Fel, Thomas, et al.
Published: (2024)
by: Fel, Thomas, et al.
Published: (2024)
Spatiotemporal Decoupling for Efficient Vision-Based Occupancy Forecasting
by: Xu, Jingyi, et al.
Published: (2024)
by: Xu, Jingyi, et al.
Published: (2024)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
by: Wu, Juncheng, et al.
Published: (2026)
by: Wu, Juncheng, et al.
Published: (2026)
Decoupled Prototype Matching with Vision Foundation Models for Few-Shot Industrial Object Detection
by: M., Hari Prasanth S., et al.
Published: (2026)
by: M., Hari Prasanth S., et al.
Published: (2026)
JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation
by: Zeng, Shuang, et al.
Published: (2025)
by: Zeng, Shuang, et al.
Published: (2025)
HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
by: Nath, Sujoy, et al.
Published: (2025)
by: Nath, Sujoy, et al.
Published: (2025)
Disease-informed Adaptation of Vision-Language Models
by: Zhang, Jiajin, et al.
Published: (2024)
by: Zhang, Jiajin, et al.
Published: (2024)
DisentangleFormer: Spatial-Channel Decoupling for Multi-Channel Vision
by: Liao, Jiashu, et al.
Published: (2025)
by: Liao, Jiashu, et al.
Published: (2025)
VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning
by: Xiao, Wenyi, et al.
Published: (2026)
by: Xiao, Wenyi, et al.
Published: (2026)
Surgical-LLaVA: Toward Surgical Scenario Understanding via Large Language and Vision Models
by: Jin, Juseong, et al.
Published: (2024)
by: Jin, Juseong, et al.
Published: (2024)
Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
by: Atabuzzaman, Md., et al.
Published: (2025)
by: Atabuzzaman, Md., et al.
Published: (2025)
Remote Sensing for Weed Detection and Control
by: Bansal, Ishita, et al.
Published: (2024)
by: Bansal, Ishita, et al.
Published: (2024)
MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence
by: Park, Woohyeon, et al.
Published: (2026)
by: Park, Woohyeon, et al.
Published: (2026)
NOVO: Unlearning-Compliant Vision Transformers
by: Roy, Soumya, et al.
Published: (2025)
by: Roy, Soumya, et al.
Published: (2025)
TinyViM: Frequency Decoupling for Tiny Hybrid Vision Mamba
by: Ma, Xiaowen, et al.
Published: (2024)
by: Ma, Xiaowen, et al.
Published: (2024)
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
by: Niu, Junbo, et al.
Published: (2025)
by: Niu, Junbo, et al.
Published: (2025)
WaveFormer: Frequency-Time Decoupled Vision Modeling with Wave Equation
by: Shu, Zishan, et al.
Published: (2026)
by: Shu, Zishan, et al.
Published: (2026)
DPL: Decoupled Prototype Learning for Enhancing Robustness of Vision-Language Transformers to Missing Modalities
by: Lu, Jueqing, et al.
Published: (2025)
by: Lu, Jueqing, et al.
Published: (2025)
Efficient Vision-Language Pre-training by Cluster Masking
by: Wei, Zihao, et al.
Published: (2024)
by: Wei, Zihao, et al.
Published: (2024)
Deep Height Decoupling for Precise Vision-based 3D Occupancy Prediction
by: Wu, Yuan, et al.
Published: (2024)
by: Wu, Yuan, et al.
Published: (2024)
Vision-Language Models for Vision Tasks: A Survey
by: Zhang, Jingyi, et al.
Published: (2023)
by: Zhang, Jingyi, et al.
Published: (2023)
VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering
by: Wang, Yanling, et al.
Published: (2025)
by: Wang, Yanling, et al.
Published: (2025)
Similar Items
-
The in-context inductive biases of vision-language models differ across modalities
by: Allen, Kelsey, et al.
Published: (2025) -
Libra: Building Decoupled Vision System on Large Language Models
by: Xu, Yifan, et al.
Published: (2024) -
Hierarchical Dual-Subspace Decoupling for Continual Learning in Vision-Language Models
by: Qin, Mengxin, et al.
Published: (2026) -
Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models
by: Ma, Kexin, et al.
Published: (2026) -
Context-Based Visual-Language Place Recognition
by: Woo, Soojin, et al.
Published: (2024)