Guardado en:
| Autores principales: | Zhao, Han, Wang, Jingbo, Song, Wenxuan, Chen, Shuai, Liu, Yang, Wang, Yan, Li, Haoang, Wang, Donglin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2602.17259 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
por: Li, Fuhao, et al.
Publicado: (2025)
por: Li, Fuhao, et al.
Publicado: (2025)
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching
por: Chen, Jiayi, et al.
Publicado: (2026)
por: Chen, Jiayi, et al.
Publicado: (2026)
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
por: Song, Wenxuan, et al.
Publicado: (2026)
por: Song, Wenxuan, et al.
Publicado: (2026)
CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding
por: Song, Wenxuan, et al.
Publicado: (2025)
por: Song, Wenxuan, et al.
Publicado: (2025)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
por: Chen, Jiayi, et al.
Publicado: (2025)
por: Chen, Jiayi, et al.
Publicado: (2025)
GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot
por: Song, Wenxuan, et al.
Publicado: (2024)
por: Song, Wenxuan, et al.
Publicado: (2024)
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
por: Song, Wenxuan, et al.
Publicado: (2026)
por: Song, Wenxuan, et al.
Publicado: (2026)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
por: Song, Wenxuan, et al.
Publicado: (2025)
por: Song, Wenxuan, et al.
Publicado: (2025)
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
por: Zhao, Han, et al.
Publicado: (2025)
por: Zhao, Han, et al.
Publicado: (2025)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
por: Song, Wenxuan, et al.
Publicado: (2025)
por: Song, Wenxuan, et al.
Publicado: (2025)
RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark
por: Lei, Huashuo, et al.
Publicado: (2026)
por: Lei, Huashuo, et al.
Publicado: (2026)
Rethinking the Practicality of Vision-language-action Model: A Comprehensive Benchmark and An Improved Baseline
por: Song, Wenxuan, et al.
Publicado: (2026)
por: Song, Wenxuan, et al.
Publicado: (2026)
VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models
por: Ge, Zirui, et al.
Publicado: (2026)
por: Ge, Zirui, et al.
Publicado: (2026)
DyGeoVLN: Infusing Dynamic Geometry Foundation Model into Vision-Language Navigation
por: Liu, Xiangchen, et al.
Publicado: (2026)
por: Liu, Xiangchen, et al.
Publicado: (2026)
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
por: Lin, Minghui, et al.
Publicado: (2025)
por: Lin, Minghui, et al.
Publicado: (2025)
RationalVLA: A Rational Vision-Language-Action Model with Dual System
por: Song, Wenxuan, et al.
Publicado: (2025)
por: Song, Wenxuan, et al.
Publicado: (2025)
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
por: Hu, Yucheng, et al.
Publicado: (2024)
por: Hu, Yucheng, et al.
Publicado: (2024)
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models
por: Zhao, Han, et al.
Publicado: (2025)
por: Zhao, Han, et al.
Publicado: (2025)
SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation
por: Huang, Jingzhi, et al.
Publicado: (2026)
por: Huang, Jingzhi, et al.
Publicado: (2026)
Dynamic Adaptive Legged Locomotion Policy via Decoupling Reaction Force Control and Gait Control
por: Wang, Renjie, et al.
Publicado: (2025)
por: Wang, Renjie, et al.
Publicado: (2025)
Scaling World Model for Hierarchical Manipulation Policies
por: Long, Qian, et al.
Publicado: (2026)
por: Long, Qian, et al.
Publicado: (2026)
Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
por: Bai, Shuanghao, et al.
Publicado: (2025)
por: Bai, Shuanghao, et al.
Publicado: (2025)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
por: Ding, Pengxiang, et al.
Publicado: (2023)
por: Ding, Pengxiang, et al.
Publicado: (2023)
X-Loco: Towards Generalist Humanoid Locomotion Control via Synergetic Policy Distillation
por: Wang, Dewei, et al.
Publicado: (2026)
por: Wang, Dewei, et al.
Publicado: (2026)
FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
por: Zhong, Zhide, et al.
Publicado: (2025)
por: Zhong, Zhide, et al.
Publicado: (2025)
Discovering Self-Protective Falling Policy for Humanoid Robot via Deep Reinforcement Learning
por: Shi, Diyuan, et al.
Publicado: (2025)
por: Shi, Diyuan, et al.
Publicado: (2025)
VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation
por: Zhong, Zhide, et al.
Publicado: (2026)
por: Zhong, Zhide, et al.
Publicado: (2026)
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
por: Liu, Yang, et al.
Publicado: (2026)
por: Liu, Yang, et al.
Publicado: (2026)
Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey
por: Bai, Shuanghao, et al.
Publicado: (2025)
por: Bai, Shuanghao, et al.
Publicado: (2025)
A Real-World Quadrupedal Locomotion Benchmark for Offline Reinforcement Learning
por: Zhang, Hongyin, et al.
Publicado: (2023)
por: Zhang, Hongyin, et al.
Publicado: (2023)
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
por: Atreya, Pranav, et al.
Publicado: (2025)
por: Atreya, Pranav, et al.
Publicado: (2025)
NFPO: Stabilized Policy Optimization of Normalizing Flow for Robotic Policy Learning
por: Shi, Diyuan, et al.
Publicado: (2026)
por: Shi, Diyuan, et al.
Publicado: (2026)
Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies
por: Wang, Yi, et al.
Publicado: (2026)
por: Wang, Yi, et al.
Publicado: (2026)
OmniGuide: Universal Guidance Fields for Enhancing Generalist Robot Policies
por: Song, Yunzhou, et al.
Publicado: (2026)
por: Song, Yunzhou, et al.
Publicado: (2026)
NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
por: Chen, Jiahong, et al.
Publicado: (2025)
por: Chen, Jiahong, et al.
Publicado: (2025)
Hardware-Free Event Cameras Temporal Synchronization Based on Event Density Alignment
por: Li, Wenxuan, et al.
Publicado: (2025)
por: Li, Wenxuan, et al.
Publicado: (2025)
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
por: Wang, Yihao, et al.
Publicado: (2025)
por: Wang, Yihao, et al.
Publicado: (2025)
Effective Tuning Strategies for Generalist Robot Manipulation Policies
por: Zhang, Wenbo, et al.
Publicado: (2024)
por: Zhang, Wenbo, et al.
Publicado: (2024)
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
por: Zhao, Wei, et al.
Publicado: (2025)
por: Zhao, Wei, et al.
Publicado: (2025)
Efficient Reinforcement Learning by Guiding Generalist World Models with Non-Curated Data
por: Zhao, Yi, et al.
Publicado: (2025)
por: Zhao, Yi, et al.
Publicado: (2025)
Ejemplares similares
-
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
por: Li, Fuhao, et al.
Publicado: (2025) -
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching
por: Chen, Jiayi, et al.
Publicado: (2026) -
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
por: Song, Wenxuan, et al.
Publicado: (2026) -
CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding
por: Song, Wenxuan, et al.
Publicado: (2025) -
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
por: Chen, Jiayi, et al.
Publicado: (2025)