Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Weiliang, Jing, Dong, Pan, Jia-Hui, Lu, Zhiwu, Liu, Yun-Hui, Li, Li Erran, Ding, Mingyu, Fu, Chi-Wing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation
by: Tang, Weiliang, et al.
Published: (2025)
by: Tang, Weiliang, et al.
Published: (2025)
Rethinking Intermediate Representation for VLM-based Robot Manipulation
by: Tang, Weiliang, et al.
Published: (2025)
by: Tang, Weiliang, et al.
Published: (2025)
Embodiment-Agnostic Action Planning via Object-Part Scene Flow
by: Tang, Weiliang, et al.
Published: (2024)
by: Tang, Weiliang, et al.
Published: (2024)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
by: Chen, Yi, et al.
Published: (2024)
by: Chen, Yi, et al.
Published: (2024)
OPA-Pack: Object-Property-Aware Robotic Bin Packing
by: Pan, Jia-Hui, et al.
Published: (2025)
by: Pan, Jia-Hui, et al.
Published: (2025)
Mixture of Horizons in Action Chunking
by: Jing, Dong, et al.
Published: (2025)
by: Jing, Dong, et al.
Published: (2025)
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval
by: Sun, Zelong, et al.
Published: (2025)
by: Sun, Zelong, et al.
Published: (2025)
MLLMRec-R1: Incentivizing Reasoning Capability in Large Language Models for Multimodal Sequential Recommendation
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
by: Huang, Wenxuan, et al.
Published: (2025)
by: Huang, Wenxuan, et al.
Published: (2025)
When would Vision-Proprioception Policies Fail in Robotic Manipulation?
by: Lu, Jingxian, et al.
Published: (2026)
by: Lu, Jingxian, et al.
Published: (2026)
Overcoming Support Dilution for Robust Few-shot Semantic Segmentation
by: Tang, Wailing, et al.
Published: (2025)
by: Tang, Wailing, et al.
Published: (2025)
VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning
by: Xu, Hengbo, et al.
Published: (2026)
by: Xu, Hengbo, et al.
Published: (2026)
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval
by: Sun, Zelong, et al.
Published: (2024)
by: Sun, Zelong, et al.
Published: (2024)
Bridging Writing Manner Gap in Visual Instruction Tuning by Creating LLM-aligned Instructions
by: Jing, Dong, et al.
Published: (2025)
by: Jing, Dong, et al.
Published: (2025)
Incentivizing Dual Process Thinking for Efficient Large Language Model Reasoning
by: Cheng, Xiaoxue, et al.
Published: (2025)
by: Cheng, Xiaoxue, et al.
Published: (2025)
From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation
by: Liu, Yibin, et al.
Published: (2026)
by: Liu, Yibin, et al.
Published: (2026)
COS3D: Collaborative Open-Vocabulary 3D Segmentation
by: Zhu, Runsong, et al.
Published: (2025)
by: Zhu, Runsong, et al.
Published: (2025)
Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models
by: Qin, Yulei, et al.
Published: (2025)
by: Qin, Yulei, et al.
Published: (2025)
C3LLM: Conditional Multimodal Content Generation Using Large Language Models
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
Current-State Opacity in Safe Partially Observed Quantum Petri Nets: True-Concurrency Semantics and Exact Symbolic Verification
by: Ding, Sichen, et al.
Published: (2026)
by: Ding, Sichen, et al.
Published: (2026)
DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
ManipulationNet: An Infrastructure for Benchmarking Real-World Robot Manipulation with Physical Skill Challenges and Embodied Multimodal Reasoning
by: Chen, Yiting, et al.
Published: (2026)
by: Chen, Yiting, et al.
Published: (2026)
Evaluating Step-by-Step Reasoning through Symbolic Verification
by: Zhang, Yi-Fan, et al.
Published: (2022)
by: Zhang, Yi-Fan, et al.
Published: (2022)
AssemLM: Spatial Reasoning Multimodal Large Language Models for Robotic Assembly
by: Jing, Zhi, et al.
Published: (2026)
by: Jing, Zhi, et al.
Published: (2026)
LMM-Incentive: Large Multimodal Model-based Incentive Design for User-Generated Content in Web 3.0
by: Wen, Jinbo, et al.
Published: (2025)
by: Wen, Jinbo, et al.
Published: (2025)
Enhancing Robotic Manipulation with AI Feedback from Multimodal Large Language Models
by: Liu, Jinyi, et al.
Published: (2024)
by: Liu, Jinyi, et al.
Published: (2024)
The application value of routine office hysteroscopy prior in patients with one embryo transferred failure and normal transvaginal ultrasound
by: Hui‐Juan Guan, et al.
Published: (2025)
by: Hui‐Juan Guan, et al.
Published: (2025)
GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
by: Lu, Guanxing, et al.
Published: (2025)
by: Lu, Guanxing, et al.
Published: (2025)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
by: Han, Xiaofeng, et al.
Published: (2025)
by: Han, Xiaofeng, et al.
Published: (2025)
PhyGrasp: Generalizing Robotic Grasping with Physics-informed Large Multimodal Models
by: Guo, Dingkun, et al.
Published: (2024)
by: Guo, Dingkun, et al.
Published: (2024)
Data‐Driven Modeling and High‐Precision Tracking Control of a Soft Continuum Manipulator: Enabling Robotic Sorting of Multiwire Cables
by: Yuan Gao, et al.
Published: (2024)
by: Yuan Gao, et al.
Published: (2024)
Multimodal Generation of Animatable 3D Human Models with AvatarForge
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
by: Fan, Linfeng, et al.
Published: (2026)
by: Fan, Linfeng, et al.
Published: (2026)
Interpretable Multimodal Misinformation Detection with Logic Reasoning
by: Liu, Hui, et al.
Published: (2023)
by: Liu, Hui, et al.
Published: (2023)
Prioritizing the Best: Incentivizing Reliable Multimodal Reasoning by Rewarding Beyond Answer Correctness
by: Jia, Mengzhao, et al.
Published: (2026)
by: Jia, Mengzhao, et al.
Published: (2026)
Amortized Reasoning Tree Search: Decoupling Proposal and Decision in Large Language Models
by: Hong, Zesheng, et al.
Published: (2026)
by: Hong, Zesheng, et al.
Published: (2026)
DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry
by: Cai, Zhenyang, et al.
Published: (2025)
by: Cai, Zhenyang, et al.
Published: (2025)
SldprtNet: A Large-Scale Multimodal Dataset for CAD Generation in Language-Driven 3D Design
by: Li, Ruogu, et al.
Published: (2026)
by: Li, Ruogu, et al.
Published: (2026)
R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
by: Yao, Huanjin, et al.
Published: (2025)
by: Yao, Huanjin, et al.
Published: (2025)
EventVL: Understand Event Streams via Multimodal Large Language Model
by: Li, Pengteng, et al.
Published: (2025)
by: Li, Pengteng, et al.
Published: (2025)
Similar Items
-
GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation
by: Tang, Weiliang, et al.
Published: (2025) -
Rethinking Intermediate Representation for VLM-based Robot Manipulation
by: Tang, Weiliang, et al.
Published: (2025) -
Embodiment-Agnostic Action Planning via Object-Part Scene Flow
by: Tang, Weiliang, et al.
Published: (2024) -
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
by: Chen, Yi, et al.
Published: (2024) -
OPA-Pack: Object-Property-Aware Robotic Bin Packing
by: Pan, Jia-Hui, et al.
Published: (2025)