CoV: Chain-of-View Prompting for Spatial Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Haoyu, Liu, Akide, Zhang, Zeyu, Wang, Weijie, Chen, Feng, Zhu, Ruihan, Haffari, Gholamreza, Zhuang, Bohan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation
by: Liu, Akide, et al.
Published: (2026)
by: Liu, Akide, et al.
Published: (2026)
ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS
by: Wang, Weijie, et al.
Published: (2025)
by: Wang, Weijie, et al.
Published: (2025)
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
by: Liu, Akide, et al.
Published: (2025)
by: Liu, Akide, et al.
Published: (2025)
Motion Mamba: Efficient and Long Sequence Motion Generation
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
Less Detail, Better Answers: Degradation-Driven Prompting for VQA
by: Han, Haoxuan, et al.
Published: (2026)
by: Han, Haoxuan, et al.
Published: (2026)
An Empirical Study on How Video-LLMs Answer Video Questions
by: Gou, Chenhui, et al.
Published: (2025)
by: Gou, Chenhui, et al.
Published: (2025)
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations
by: Ke, Fucai, et al.
Published: (2026)
by: Ke, Fucai, et al.
Published: (2026)
KMM: Key Frame Mask Mamba for Extended Motion Generation
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models
by: Shiri, Fatemeh, et al.
Published: (2024)
by: Shiri, Fatemeh, et al.
Published: (2024)
FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation
by: Zhou, Junkang, et al.
Published: (2026)
by: Zhou, Junkang, et al.
Published: (2026)
VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction
by: Wang, Weijie, et al.
Published: (2025)
by: Wang, Weijie, et al.
Published: (2025)
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
by: Tao, Xingjian, et al.
Published: (2026)
by: Tao, Xingjian, et al.
Published: (2026)
R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs
by: Xie, Jiahao, et al.
Published: (2026)
by: Xie, Jiahao, et al.
Published: (2026)
ESA: Annotation-Efficient Active Learning for Semantic Segmentation
by: Ge, Jinchao, et al.
Published: (2024)
by: Ge, Jinchao, et al.
Published: (2024)
Revisiting Depth Representations for Feed-Forward 3D Gaussian Splatting
by: Shi, Duochao, et al.
Published: (2025)
by: Shi, Duochao, et al.
Published: (2025)
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
by: Cheng, An-Chieh, et al.
Published: (2024)
by: Cheng, An-Chieh, et al.
Published: (2024)
Physics-Grounded Motion Forecasting via Equation Discovery for Trajectory-Guided Image-to-Video Generation
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction
by: Wang, Weijie, et al.
Published: (2026)
by: Wang, Weijie, et al.
Published: (2026)
Geometrically-Constrained Agent for Spatial Reasoning
by: Chen, Zeren, et al.
Published: (2025)
by: Chen, Zeren, et al.
Published: (2025)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
by: Feng, Zhiyuan, et al.
Published: (2025)
by: Feng, Zhiyuan, et al.
Published: (2025)
SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning
by: Liu, Yuecheng, et al.
Published: (2025)
by: Liu, Yuecheng, et al.
Published: (2025)
Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought
by: Peng, Yi, et al.
Published: (2025)
by: Peng, Yi, et al.
Published: (2025)
Spatial Chain-of-Thought: Bridging Understanding and Generation Models for Spatial Reasoning Generation
by: Chen, Wei, et al.
Published: (2026)
by: Chen, Wei, et al.
Published: (2026)
CoT-Pose: Chain-of-Thought Reasoning for 3D Pose Generation from Abstract Prompts
by: Cha, Junuk, et al.
Published: (2025)
by: Cha, Junuk, et al.
Published: (2025)
LongVLM: Efficient Long Video Understanding via Large Language Models
by: Weng, Yuetian, et al.
Published: (2024)
by: Weng, Yuetian, et al.
Published: (2024)
Streaming Video Diffusion: Online Video Editing with Diffusion Models
by: Chen, Feng, et al.
Published: (2024)
by: Chen, Feng, et al.
Published: (2024)
Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning
by: Sinha, Sanchit, et al.
Published: (2026)
by: Sinha, Sanchit, et al.
Published: (2026)
InterCoG: Towards Spatially Precise Image Editing with Interleaved Chain-of-Grounding Reasoning
by: Wan, Yecong, et al.
Published: (2026)
by: Wan, Yecong, et al.
Published: (2026)
Learning Multi-View Spatial Reasoning from Cross-View Relations
by: Jeong, Suchae, et al.
Published: (2026)
by: Jeong, Suchae, et al.
Published: (2026)
Seeing Through the Chain: Mitigate Hallucination in Multimodal Reasoning Models via CoT Compression and Contrastive Preference Optimization
by: Fang, Hao, et al.
Published: (2026)
by: Fang, Hao, et al.
Published: (2026)
CoS: Chain-of-Shot Prompting for Long Video Understanding
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
by: Lu, Zhenyu, et al.
Published: (2025)
by: Lu, Zhenyu, et al.
Published: (2025)
MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance
by: Bai, Xuehai, et al.
Published: (2026)
by: Bai, Xuehai, et al.
Published: (2026)
CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs
by: Zhang, Daoan, et al.
Published: (2024)
by: Zhang, Daoan, et al.
Published: (2024)
MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse Views
by: Chen, Yuedong, et al.
Published: (2024)
by: Chen, Yuedong, et al.
Published: (2024)
CoDA: Instructive Chain-of-Domain Adaptation with Severity-Aware Visual Prompt Tuning
by: Gong, Ziyang, et al.
Published: (2024)
by: Gong, Ziyang, et al.
Published: (2024)
Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
by: Li, Qiming, et al.
Published: (2025)
by: Li, Qiming, et al.
Published: (2025)
Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models
by: Xu, Qinwu, et al.
Published: (2026)
by: Xu, Qinwu, et al.
Published: (2026)
Similar Items
-
ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation
by: Liu, Akide, et al.
Published: (2026) -
ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS
by: Wang, Weijie, et al.
Published: (2025) -
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
by: Liu, Akide, et al.
Published: (2025) -
Motion Mamba: Efficient and Long Sequence Motion Generation
by: Zhang, Zeyu, et al.
Published: (2024) -
InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation
by: Zhang, Zeyu, et al.
Published: (2024)