CVT-Bench: Counterfactual Viewpoint Transformations Reveal Unstable Spatial Representations in Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Vellamcheti, Shanmukha, Kothapalli, Uday Kiran, Bhowmick, Disharee, Aakur, Sathyanarayanan N. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Probabilistic Jump-Diffusion Framework for Open-World Egocentric Activity Recognition
by: Kundu, Sanjoy, et al.
Published: (2025)
by: Kundu, Sanjoy, et al.
Published: (2025)
Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection
by: Vellamcheti, Shanmukha, et al.
Published: (2025)
by: Vellamcheti, Shanmukha, et al.
Published: (2025)
ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity Recognition
by: Kundu, Sanjoy, et al.
Published: (2025)
by: Kundu, Sanjoy, et al.
Published: (2025)
STaTS: Structure-Aware Temporal Sequence Summarization via Statistical Window Merging
by: Bhowmick, Disharee, et al.
Published: (2025)
by: Bhowmick, Disharee, et al.
Published: (2025)
Self-supervised Multi-actor Social Activity Understanding in Streaming Videos
by: Trehan, Shubham, et al.
Published: (2024)
by: Trehan, Shubham, et al.
Published: (2024)
CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
Discovering Novel Actions from Open World Egocentric Videos with Object-Grounded Visual Commonsense Reasoning
by: Kundu, Sanjoy, et al.
Published: (2023)
by: Kundu, Sanjoy, et al.
Published: (2023)
Generalized Event Partonomy Inference with Structured Hierarchical Predictive Learning
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
ALGO: Object-Grounded Visual Commonsense Reasoning for Open-World Egocentric Action Recognition
by: Kundu, Sanjoy, et al.
Published: (2024)
by: Kundu, Sanjoy, et al.
Published: (2024)
Capturing Temporal Components for Time Series Classification
by: Vavilthota, Venkata Ragavendra, et al.
Published: (2024)
by: Vavilthota, Venkata Ragavendra, et al.
Published: (2024)
EASE: Embodied Active Event Perception via Self-Supervised Energy Minimization
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
FSP-DETR: Few-Shot Prototypical Parasitic Ova Detection
by: Trehan, Shubham, et al.
Published: (2025)
by: Trehan, Shubham, et al.
Published: (2025)
Embedding Textual Information in Images Using Quinary Pixel Combinations
by: Kandala, A V Uday Kiran
Published: (2026)
by: Kandala, A V Uday Kiran
Published: (2026)
CVT-xRF: Contrastive In-Voxel Transformer for 3D Consistent Radiance Fields from Sparse Inputs
by: Zhong, Yingji, et al.
Published: (2024)
by: Zhong, Yingji, et al.
Published: (2024)
Improving Viewpoint-Independent Object-Centric Representations through Active Viewpoint Selection
by: Huang, Yinxuan, et al.
Published: (2024)
by: Huang, Yinxuan, et al.
Published: (2024)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
by: Xu, Zelin, et al.
Published: (2026)
by: Xu, Zelin, et al.
Published: (2026)
SpineBench: Benchmarking Multimodal LLMs for Spinal Pathology Analysis
by: Zhang, Chenghanyu, et al.
Published: (2025)
by: Zhang, Chenghanyu, et al.
Published: (2025)
3D Segmentation Using Viewpoint-Dependent Spatial Relationships
by: Nanri, Ayaka, et al.
Published: (2026)
by: Nanri, Ayaka, et al.
Published: (2026)
HINT: Learning Complete Human Neural Representations from Limited Viewpoints
by: Sanvito, Alessandro, et al.
Published: (2024)
by: Sanvito, Alessandro, et al.
Published: (2024)
CVT-Occ: Cost Volume Temporal Fusion for 3D Occupancy Prediction
by: Ye, Zhangchen, et al.
Published: (2024)
by: Ye, Zhangchen, et al.
Published: (2024)
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
by: Qian, Yusu, et al.
Published: (2024)
by: Qian, Yusu, et al.
Published: (2024)
SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
by: Guo, Jiajie, et al.
Published: (2025)
by: Guo, Jiajie, et al.
Published: (2025)
A Framework for Fluid Motion Estimation using a Constraint-Based Refinement Approach
by: Doshi, Hirak, et al.
Published: (2020)
by: Doshi, Hirak, et al.
Published: (2020)
E3VS-Bench: A Benchmark for Viewpoint-Dependent Active Perception in 3D Gaussian Splatting Scenes
by: Sakamoto, Koya, et al.
Published: (2026)
by: Sakamoto, Koya, et al.
Published: (2026)
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
by: Guo, Zichun, et al.
Published: (2026)
by: Guo, Zichun, et al.
Published: (2026)
Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models
by: Rezaei, Parham, et al.
Published: (2025)
by: Rezaei, Parham, et al.
Published: (2025)
StreamSTGS: Streaming Spatial and Temporal Gaussian Grids for Real-Time Free-Viewpoint Video
by: Ke, Zhihui, et al.
Published: (2025)
by: Ke, Zhihui, et al.
Published: (2025)
ViewpointDepth: A New Dataset for Monocular Depth Estimation Under Viewpoint Shifts
by: Pjetri, Aurel, et al.
Published: (2024)
by: Pjetri, Aurel, et al.
Published: (2024)
Evo-Retriever: LLM-Guided Curriculum Evolution with Viewpoint-Pathway Collaboration for Multimodal Document Retrieval
by: Li, Weiqing, et al.
Published: (2026)
by: Li, Weiqing, et al.
Published: (2026)
Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving
by: Zhou, Hao, et al.
Published: (2024)
by: Zhou, Hao, et al.
Published: (2024)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
by: Ma, David, et al.
Published: (2025)
by: Ma, David, et al.
Published: (2025)
STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
by: Li, Zongzhao, et al.
Published: (2025)
by: Li, Zongzhao, et al.
Published: (2025)
From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs
by: Zhang, Le, et al.
Published: (2026)
by: Zhang, Le, et al.
Published: (2026)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
SpatialBench: Is Your Spatial Foundation Model an All-Round Player?
by: Peng, Haosong, et al.
Published: (2026)
by: Peng, Haosong, et al.
Published: (2026)
CLOVER: Context-aware Long-term Object Viewpoint- and Environment- Invariant Representation Learning
by: Lee, Dongmyeong, et al.
Published: (2024)
by: Lee, Dongmyeong, et al.
Published: (2024)
XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?
by: Wang, Fengxiang, et al.
Published: (2025)
by: Wang, Fengxiang, et al.
Published: (2025)
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
by: Ye, Wenqian, et al.
Published: (2024)
by: Ye, Wenqian, et al.
Published: (2024)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
DSI-Bench: A Benchmark for Dynamic Spatial Intelligence
by: Zhang, Ziang, et al.
Published: (2025)
by: Zhang, Ziang, et al.
Published: (2025)
Similar Items
-
A Probabilistic Jump-Diffusion Framework for Open-World Egocentric Activity Recognition
by: Kundu, Sanjoy, et al.
Published: (2025) -
Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection
by: Vellamcheti, Shanmukha, et al.
Published: (2025) -
ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity Recognition
by: Kundu, Sanjoy, et al.
Published: (2025) -
STaTS: Structure-Aware Temporal Sequence Summarization via Statistical Window Merging
by: Bhowmick, Disharee, et al.
Published: (2025) -
Self-supervised Multi-actor Social Activity Understanding in Streaming Videos
by: Trehan, Shubham, et al.
Published: (2024)