Rethinking Intermediate Representation for VLM-based Robot Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Weiliang, Gao, Jialin, Pan, Jia-Hui, Wang, Gang, Li, Li Erran, Liu, Yunhui, Ding, Mingyu, Heng, Pheng-Ann, Fu, Chi-Wing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation
by: Tang, Weiliang, et al.
Published: (2025)
by: Tang, Weiliang, et al.
Published: (2025)
Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation
by: Tang, Weiliang, et al.
Published: (2025)
by: Tang, Weiliang, et al.
Published: (2025)
OPA-Pack: Object-Property-Aware Robotic Bin Packing
by: Pan, Jia-Hui, et al.
Published: (2025)
by: Pan, Jia-Hui, et al.
Published: (2025)
Overcoming Support Dilution for Robust Few-shot Semantic Segmentation
by: Tang, Wailing, et al.
Published: (2025)
by: Tang, Wailing, et al.
Published: (2025)
COS3D: Collaborative Open-Vocabulary 3D Segmentation
by: Zhu, Runsong, et al.
Published: (2025)
by: Zhu, Runsong, et al.
Published: (2025)
Rethinking End-to-End 2D to 3D Scene Segmentation in Gaussian Splatting
by: Zhu, Runsong, et al.
Published: (2025)
by: Zhu, Runsong, et al.
Published: (2025)
UniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose Estimation
by: Wang, Yinqiao, et al.
Published: (2025)
by: Wang, Yinqiao, et al.
Published: (2025)
SiMA-Hand: Boosting 3D Hand-Mesh Reconstruction by Single-to-Multi-View Adaptation
by: Wang, Yinqiao, et al.
Published: (2024)
by: Wang, Yinqiao, et al.
Published: (2024)
DisCo-Layout: Disentangling and Coordinating Semantic and Physical Refinement in a Multi-Agent Framework for 3D Indoor Layout Synthesis
by: Gao, Jialin, et al.
Published: (2025)
by: Gao, Jialin, et al.
Published: (2025)
Embodiment-Agnostic Action Planning via Object-Part Scene Flow
by: Tang, Weiliang, et al.
Published: (2024)
by: Tang, Weiliang, et al.
Published: (2024)
Unveiling Deep Shadows: A Survey and Benchmark on Image and Video Shadow Detection, Removal, and Generation in the Deep Learning Era
by: Hu, Xiaowei, et al.
Published: (2024)
by: Hu, Xiaowei, et al.
Published: (2024)
PCF-Lift: Panoptic Lifting by Probabilistic Contrastive Fusion
by: Zhu, Runsong, et al.
Published: (2024)
by: Zhu, Runsong, et al.
Published: (2024)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
by: Chen, Yi, et al.
Published: (2024)
by: Chen, Yi, et al.
Published: (2024)
Video Instance Shadow Detection Under the Sun and Sky
by: Xing, Zhenghao, et al.
Published: (2022)
by: Xing, Zhenghao, et al.
Published: (2022)
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
by: Xing, Zhenghao, et al.
Published: (2025)
by: Xing, Zhenghao, et al.
Published: (2025)
Towards Real-World Adverse Weather Image Restoration: Enhancing Clearness and Semantics with Vision-Language Models
by: Xu, Jiaqi, et al.
Published: (2024)
by: Xu, Jiaqi, et al.
Published: (2024)
Revisiting Shadow Detection: A New Benchmark Dataset for Complex World
by: Hu, Xiaowei, et al.
Published: (2019)
by: Hu, Xiaowei, et al.
Published: (2019)
Rethinking VLM Representation for VLA Initialization
by: Lin, Weifeng, et al.
Published: (2026)
by: Lin, Weifeng, et al.
Published: (2026)
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
by: Nasiriany, Soroush, et al.
Published: (2024)
by: Nasiriany, Soroush, et al.
Published: (2024)
LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning
by: Chen, Hao, et al.
Published: (2026)
by: Chen, Hao, et al.
Published: (2026)
Hand-Shadow Poser
by: Xu, Hao, et al.
Published: (2025)
by: Xu, Hao, et al.
Published: (2025)
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Perceiving, Reasoning, Adapting: A Dual-Layer Framework for VLM-Guided Precision Robotic Manipulation
by: Jia, Qingxuan, et al.
Published: (2025)
by: Jia, Qingxuan, et al.
Published: (2025)
RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models
by: Yan, Qiao, et al.
Published: (2025)
by: Yan, Qiao, et al.
Published: (2025)
IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation
by: Zhou, Donghao, et al.
Published: (2025)
by: Zhou, Donghao, et al.
Published: (2025)
Robot Collapse: Supply Chain Backdoor Attacks Against VLM-based Robotic Manipulation
by: Wang, Xianlong, et al.
Published: (2024)
by: Wang, Xianlong, et al.
Published: (2024)
A Collaborative Extended Reality Prototype for 3D Surgical Planning and Visualization
by: Qiu, Shi, et al.
Published: (2026)
by: Qiu, Shi, et al.
Published: (2026)
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
by: Gu, Chenyang, et al.
Published: (2025)
by: Gu, Chenyang, et al.
Published: (2025)
RiboSphere: Learning Unified and Efficient Representations of RNA Structures
by: Zhang, Zhou, et al.
Published: (2026)
by: Zhang, Zhou, et al.
Published: (2026)
Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision Making
by: Wang, Yihan, et al.
Published: (2025)
by: Wang, Yihan, et al.
Published: (2025)
Coordinated 2D-3D Visualization of Volumetric Medical Data in XR with Multimodal Interactions
by: Liu, Qixuan, et al.
Published: (2025)
by: Liu, Qixuan, et al.
Published: (2025)
EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning
by: Xie, Binzhu, et al.
Published: (2026)
by: Xie, Binzhu, et al.
Published: (2026)
CvhSlicer 2.0: Immersive and Interactive Visualization of Chinese Visible Human Data in XR Environments
by: Qiu, Yue, et al.
Published: (2025)
by: Qiu, Yue, et al.
Published: (2025)
Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt Framework
by: Zhu, Yu, et al.
Published: (2026)
by: Zhu, Yu, et al.
Published: (2026)
Mixture of Horizons in Action Chunking
by: Jing, Dong, et al.
Published: (2025)
by: Jing, Dong, et al.
Published: (2025)
Score the Steps, Not Just the Goal: VLM-Based Subgoal Evaluation for Robotic Manipulation
by: ElMallah, Ramy, et al.
Published: (2025)
by: ElMallah, Ramy, et al.
Published: (2025)
Improving AlphaFlow for Efficient Protein Ensembles Generation
by: Li, Shaoning, et al.
Published: (2024)
by: Li, Shaoning, et al.
Published: (2024)
ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation
by: Feng, Youhe, et al.
Published: (2026)
by: Feng, Youhe, et al.
Published: (2026)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Similar Items
-
GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation
by: Tang, Weiliang, et al.
Published: (2025) -
Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation
by: Tang, Weiliang, et al.
Published: (2025) -
OPA-Pack: Object-Property-Aware Robotic Bin Packing
by: Pan, Jia-Hui, et al.
Published: (2025) -
Overcoming Support Dilution for Robust Few-shot Semantic Segmentation
by: Tang, Wailing, et al.
Published: (2025) -
COS3D: Collaborative Open-Vocabulary 3D Segmentation
by: Zhu, Runsong, et al.
Published: (2025)