Salvato in:
| Autori principali: | Song, Wenxuan, Chen, Jiayi, Sun, Xiaoquan, Lei, Huashuo, Qin, Yikai, Zhao, Wei, Ding, Pengxiang, Zhao, Han, Wang, Tongxin, Hou, Pengxu, Zhong, Zhide, Yan, Haodong, Wang, Donglin, Ma, Jun, Li, Haoang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2602.22663 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark
di: Lei, Huashuo, et al.
Pubblicazione: (2026)
di: Lei, Huashuo, et al.
Pubblicazione: (2026)
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
di: Li, Fuhao, et al.
Pubblicazione: (2025)
di: Li, Fuhao, et al.
Pubblicazione: (2025)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
di: Chen, Jiayi, et al.
Pubblicazione: (2025)
di: Chen, Jiayi, et al.
Pubblicazione: (2025)
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
di: Song, Wenxuan, et al.
Pubblicazione: (2026)
di: Song, Wenxuan, et al.
Pubblicazione: (2026)
FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
di: Zhong, Zhide, et al.
Pubblicazione: (2025)
di: Zhong, Zhide, et al.
Pubblicazione: (2025)
VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation
di: Zhong, Zhide, et al.
Pubblicazione: (2026)
di: Zhong, Zhide, et al.
Pubblicazione: (2026)
Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
di: Song, Wenxuan, et al.
Pubblicazione: (2026)
di: Song, Wenxuan, et al.
Pubblicazione: (2026)
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
di: Zhao, Han, et al.
Pubblicazione: (2025)
di: Zhao, Han, et al.
Pubblicazione: (2025)
Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
di: Wang, Yihao, et al.
Pubblicazione: (2025)
di: Wang, Yihao, et al.
Pubblicazione: (2025)
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models
di: Zhao, Han, et al.
Pubblicazione: (2025)
di: Zhao, Han, et al.
Pubblicazione: (2025)
Dream-SLAM: Dreaming the Unseen for Active SLAM in Dynamic Environments
di: Meng, Xiangqi, et al.
Pubblicazione: (2026)
di: Meng, Xiangqi, et al.
Pubblicazione: (2026)
RationalVLA: A Rational Vision-Language-Action Model with Dual System
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
DyGeoVLN: Infusing Dynamic Geometry Foundation Model into Vision-Language Navigation
di: Liu, Xiangchen, et al.
Pubblicazione: (2026)
di: Liu, Xiangchen, et al.
Pubblicazione: (2026)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
di: Ding, Pengxiang, et al.
Pubblicazione: (2023)
di: Ding, Pengxiang, et al.
Pubblicazione: (2023)
Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation
di: Yan, Haodong, et al.
Pubblicazione: (2025)
di: Yan, Haodong, et al.
Pubblicazione: (2025)
FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment
di: Zhao, Han, et al.
Pubblicazione: (2026)
di: Zhao, Han, et al.
Pubblicazione: (2026)
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
di: Zhao, Wei, et al.
Pubblicazione: (2025)
di: Zhao, Wei, et al.
Pubblicazione: (2025)
Rethinking Target Label Conditioning in Adversarial Attacks: A 2D Tensor-Guided Generative Approach
di: Liu, Hangyu, et al.
Pubblicazione: (2025)
di: Liu, Hangyu, et al.
Pubblicazione: (2025)
DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models
di: Zhong, Zhide, et al.
Pubblicazione: (2026)
di: Zhong, Zhide, et al.
Pubblicazione: (2026)
S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
di: Yan, Haodong, et al.
Pubblicazione: (2026)
di: Yan, Haodong, et al.
Pubblicazione: (2026)
GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot
di: Song, Wenxuan, et al.
Pubblicazione: (2024)
di: Song, Wenxuan, et al.
Pubblicazione: (2024)
HCSG: Human-Centric Semantic-Geometric Reasoning for Vision-Language Navigation
di: Xu, Haoxuan, et al.
Pubblicazione: (2026)
di: Xu, Haoxuan, et al.
Pubblicazione: (2026)
VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
di: Zhao, Wei, et al.
Pubblicazione: (2025)
di: Zhao, Wei, et al.
Pubblicazione: (2025)
Biological Characteristics of Exosomes and Their Applications in Aquatic Organisms
di: Huashuo Tang, et al.
Pubblicazione: (2026)
di: Huashuo Tang, et al.
Pubblicazione: (2026)
ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
di: Cui, Can, et al.
Pubblicazione: (2024)
di: Cui, Can, et al.
Pubblicazione: (2024)
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching
di: Chen, Jiayi, et al.
Pubblicazione: (2026)
di: Chen, Jiayi, et al.
Pubblicazione: (2026)
Insights into Microstructural Evolution and Properties Enhancement in Cu‐15Ni‐8Sn Strips via High‐Density Dislocations, Nanotwins, and Nanoprecipitates
di: Zhibao Xie, et al.
Pubblicazione: (2025)
di: Zhibao Xie, et al.
Pubblicazione: (2025)
Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference
di: Zhao, Han, et al.
Pubblicazione: (2024)
di: Zhao, Han, et al.
Pubblicazione: (2024)
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
di: Liu, Yang, et al.
Pubblicazione: (2026)
di: Liu, Yang, et al.
Pubblicazione: (2026)
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
di: Lin, Minghui, et al.
Pubblicazione: (2025)
di: Lin, Minghui, et al.
Pubblicazione: (2025)
SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
di: Liu, Yang, et al.
Pubblicazione: (2025)
di: Liu, Yang, et al.
Pubblicazione: (2025)
ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning
di: Zhang, Hongyin, et al.
Pubblicazione: (2025)
di: Zhang, Hongyin, et al.
Pubblicazione: (2025)
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
di: Liu, Yang, et al.
Pubblicazione: (2024)
di: Liu, Yang, et al.
Pubblicazione: (2024)
Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal Transport
di: Sun, Mingyang, et al.
Pubblicazione: (2025)
di: Sun, Mingyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark
di: Lei, Huashuo, et al.
Pubblicazione: (2026) -
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
di: Li, Fuhao, et al.
Pubblicazione: (2025) -
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
di: Song, Wenxuan, et al.
Pubblicazione: (2025) -
CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding
di: Song, Wenxuan, et al.
Pubblicazione: (2025) -
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
di: Song, Wenxuan, et al.
Pubblicazione: (2025)