SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yang, Ma, Ming, Yu, Xiaomin, Ding, Pengxiang, Zhao, Han, Sun, Mingyang, Huang, Siteng, Wang, Donglin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference
by: Zhao, Han, et al.
Published: (2024)
by: Zhao, Han, et al.
Published: (2024)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
by: Ding, Pengxiang, et al.
Published: (2023)
by: Ding, Pengxiang, et al.
Published: (2023)
CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
by: Gong, Zhefei, et al.
Published: (2024)
by: Gong, Zhefei, et al.
Published: (2024)
Score and Distribution Matching Policy: Advanced Accelerated Visuomotor Policies via Matched Distillation
by: Jia, Bofang, et al.
Published: (2024)
by: Jia, Bofang, et al.
Published: (2024)
ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
by: Cui, Can, et al.
Published: (2024)
by: Cui, Can, et al.
Published: (2024)
Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal Transport
by: Sun, Mingyang, et al.
Published: (2025)
by: Sun, Mingyang, et al.
Published: (2025)
Iterative Refinement of Flow Policies in Probability Space for Online Reinforcement Learning
by: Sun, Mingyang, et al.
Published: (2025)
by: Sun, Mingyang, et al.
Published: (2025)
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
by: Zhao, Han, et al.
Published: (2025)
by: Zhao, Han, et al.
Published: (2025)
QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning
by: Tong, Xinyang, et al.
Published: (2024)
by: Tong, Xinyang, et al.
Published: (2024)
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
by: Zhao, Wei, et al.
Published: (2025)
by: Zhao, Wei, et al.
Published: (2025)
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
by: Lin, Minghui, et al.
Published: (2025)
by: Lin, Minghui, et al.
Published: (2025)
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
by: Li, Fuhao, et al.
Published: (2025)
by: Li, Fuhao, et al.
Published: (2025)
CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
by: Zhao, Wei, et al.
Published: (2025)
by: Zhao, Wei, et al.
Published: (2025)
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
by: Yuan, Tianyuan, et al.
Published: (2025)
by: Yuan, Tianyuan, et al.
Published: (2025)
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
by: Fan, Yiguo, et al.
Published: (2025)
by: Fan, Yiguo, et al.
Published: (2025)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
by: Chen, Jiayi, et al.
Published: (2025)
by: Chen, Jiayi, et al.
Published: (2025)
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models
by: Zhao, Han, et al.
Published: (2025)
by: Zhao, Han, et al.
Published: (2025)
Robust Online Residual Refinement via Koopman-Guided Dynamics Modeling
by: Gong, Zhefei, et al.
Published: (2025)
by: Gong, Zhefei, et al.
Published: (2025)
Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
by: Lian, Shijie, et al.
Published: (2025)
by: Lian, Shijie, et al.
Published: (2025)
Enhancing Adversarial Transferability via Component-Wise Transformation
by: Liu, Hangyu, et al.
Published: (2025)
by: Liu, Hangyu, et al.
Published: (2025)
Rethinking Target Label Conditioning in Adversarial Attacks: A 2D Tensor-Guided Generative Approach
by: Liu, Hangyu, et al.
Published: (2025)
by: Liu, Hangyu, et al.
Published: (2025)
Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration
by: Han, Yuhang, et al.
Published: (2024)
by: Han, Yuhang, et al.
Published: (2024)
CUBic: Coordinated Unified Bimanual Perception and Control Framework
by: Wang, Xingyu, et al.
Published: (2026)
by: Wang, Xingyu, et al.
Published: (2026)
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
by: Wang, Yihao, et al.
Published: (2025)
by: Wang, Yihao, et al.
Published: (2025)
DepthCache: Depth-Guided Training-Free Visual Token Merging for Vision-Language-Action Model Inference
by: Li, Yuquan, et al.
Published: (2026)
by: Li, Yuquan, et al.
Published: (2026)
VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
by: Li, Hengtao, et al.
Published: (2025)
by: Li, Hengtao, et al.
Published: (2025)
ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning
by: Zhang, Hongyin, et al.
Published: (2025)
by: Zhang, Hongyin, et al.
Published: (2025)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
Can Vision-Language Models Think from the Sky? Unifying UAV Reasoning and Generation
by: Sun, Jintao, et al.
Published: (2026)
by: Sun, Jintao, et al.
Published: (2026)
Perceptio: Perception Enhanced Vision Language Models via Spatial Token Generation
by: Li, Yuchen, et al.
Published: (2026)
by: Li, Yuchen, et al.
Published: (2026)
Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models
by: Zeng, Yu, et al.
Published: (2025)
by: Zeng, Yu, et al.
Published: (2025)
SSR: Pushing the Limit of Spatial Intelligence with Structured Scene Reasoning
by: Zhang, Yi, et al.
Published: (2026)
by: Zhang, Yi, et al.
Published: (2026)
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
by: Song, Wenxuan, et al.
Published: (2026)
by: Song, Wenxuan, et al.
Published: (2026)
Exploring Spatial Representation to Enhance LLM Reasoning in Aerial Vision-Language Navigation
by: Gao, Yunpeng, et al.
Published: (2024)
by: Gao, Yunpeng, et al.
Published: (2024)
Masked Depth Modeling for Spatial Perception
by: Tan, Bin, et al.
Published: (2026)
by: Tan, Bin, et al.
Published: (2026)
Evaluation and LLM-Guided Learning of ICD Coding Rationales
by: Li, Mingyang, et al.
Published: (2025)
by: Li, Mingyang, et al.
Published: (2025)
Similar Items
-
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
by: Liu, Yang, et al.
Published: (2024) -
Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference
by: Zhao, Han, et al.
Published: (2024) -
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
by: Ding, Pengxiang, et al.
Published: (2023) -
CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
by: Gong, Zhefei, et al.
Published: (2024) -
Score and Distribution Matching Policy: Advanced Accelerated Visuomotor Policies via Matched Distillation
by: Jia, Bofang, et al.
Published: (2024)