VLANeXt: Recipes for Building Strong VLA Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Xiao-Ming, Fan, Bin, Liao, Kang, Jiang, Jian-Jian, Yang, Runze, Luo, Yihang, Wu, Zhonghua, Zheng, Wei-Shi, Loy, Chen Change |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DHAGrasp: Synthesizing Affordance-Aware Dual-Hand Grasps with Text Instructions
di: Li, Quanzhou, et al.
Pubblicazione: (2025)
di: Li, Quanzhou, et al.
Pubblicazione: (2025)
MOWA: Multiple-in-One Image Warping Model
di: Liao, Kang, et al.
Pubblicazione: (2024)
di: Liao, Kang, et al.
Pubblicazione: (2024)
Decomposed Object Manipulation via Dual-Actor Policy
di: Fan, Bin, et al.
Pubblicazione: (2025)
di: Fan, Bin, et al.
Pubblicazione: (2025)
OHP-RL: Online Human Preference as Guidance in Reinforcement Learning for Robot Manipulation
di: Mo, Yunyang, et al.
Pubblicazione: (2026)
di: Mo, Yunyang, et al.
Pubblicazione: (2026)
Real-to-Sim Grasp: Rethinking the Gap between Simulation and Real World in Grasp Detection
di: Cai, Jia-Feng, et al.
Pubblicazione: (2024)
di: Cai, Jia-Feng, et al.
Pubblicazione: (2024)
AffordDexGrasp: Open-set Language-guided Dexterous Grasp with Generalizable-Instructive Affordance
di: Wei, Yi-Lin, et al.
Pubblicazione: (2025)
di: Wei, Yi-Lin, et al.
Pubblicazione: (2025)
An Economic Framework for 6-DoF Grasp Detection
di: Wu, Xiao-Ming, et al.
Pubblicazione: (2024)
di: Wu, Xiao-Ming, et al.
Pubblicazione: (2024)
SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models
di: Li, Meng, et al.
Pubblicazione: (2025)
di: Li, Meng, et al.
Pubblicazione: (2025)
Grasp as You Say: Language-guided Dexterous Grasp Generation
di: Wei, Yi-Lin, et al.
Pubblicazione: (2024)
di: Wei, Yi-Lin, et al.
Pubblicazione: (2024)
Rethinking Bimanual Robotic Manipulation: Learning with Decoupled Interaction Framework
di: Jiang, Jian-Jian, et al.
Pubblicazione: (2025)
di: Jiang, Jian-Jian, et al.
Pubblicazione: (2025)
VLA Model-Expert Collaboration for Bi-directional Manipulation Learning
di: Xiang, Tian-Yu, et al.
Pubblicazione: (2025)
di: Xiang, Tian-Yu, et al.
Pubblicazione: (2025)
StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early Observation
di: Shi, Yiran, et al.
Pubblicazione: (2026)
di: Shi, Yiran, et al.
Pubblicazione: (2026)
Next Visual Granularity Generation
di: Wang, Yikai, et al.
Pubblicazione: (2025)
di: Wang, Yikai, et al.
Pubblicazione: (2025)
Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance
di: Wang, Runze, et al.
Pubblicazione: (2026)
di: Wang, Runze, et al.
Pubblicazione: (2026)
Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
di: Fan, Cunxin, et al.
Pubblicazione: (2025)
di: Fan, Cunxin, et al.
Pubblicazione: (2025)
How Fast Can I Run My VLA? Demystifying VLA Inference Performance with VLA-Perf
di: Jiang, Wenqi, et al.
Pubblicazione: (2026)
di: Jiang, Wenqi, et al.
Pubblicazione: (2026)
PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models
di: Guo, Xinyu, et al.
Pubblicazione: (2026)
di: Guo, Xinyu, et al.
Pubblicazione: (2026)
RoboGene: Boosting VLA Pre-training via Diversity-Driven Agentic Framework for Real-World Task Generation
di: Zhang, Yixue, et al.
Pubblicazione: (2026)
di: Zhang, Yixue, et al.
Pubblicazione: (2026)
VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
di: Zhou, Hui, et al.
Pubblicazione: (2025)
di: Zhou, Hui, et al.
Pubblicazione: (2025)
GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
di: Sun, Lin, et al.
Pubblicazione: (2025)
di: Sun, Lin, et al.
Pubblicazione: (2025)
UGotMe: An Embodied System for Affective Human-Robot Interaction
di: Li, Peizhen, et al.
Pubblicazione: (2024)
di: Li, Peizhen, et al.
Pubblicazione: (2024)
World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training
di: Xiao, Junjin, et al.
Pubblicazione: (2025)
di: Xiao, Junjin, et al.
Pubblicazione: (2025)
NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models
di: Zhu, Ziyue, et al.
Pubblicazione: (2026)
di: Zhu, Ziyue, et al.
Pubblicazione: (2026)
A Feasibility-Enhanced Control Barrier Function Method for Multi-UAV Collision Avoidance
di: Zhong, Qishen, et al.
Pubblicazione: (2026)
di: Zhong, Qishen, et al.
Pubblicazione: (2026)
VLA-0: Building State-of-the-Art VLAs with Zero Modification
di: Goyal, Ankit, et al.
Pubblicazione: (2025)
di: Goyal, Ankit, et al.
Pubblicazione: (2025)
RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models
di: Zhang, Hongyin, et al.
Pubblicazione: (2025)
di: Zhang, Hongyin, et al.
Pubblicazione: (2025)
SimVLA: A Simple VLA Baseline for Robotic Manipulation
di: Luo, Yuankai, et al.
Pubblicazione: (2026)
di: Luo, Yuankai, et al.
Pubblicazione: (2026)
Arbitrary-steps Image Super-resolution via Diffusion Inversion
di: Yue, Zongsheng, et al.
Pubblicazione: (2024)
di: Yue, Zongsheng, et al.
Pubblicazione: (2024)
Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization
di: Huang, Jialei, et al.
Pubblicazione: (2025)
di: Huang, Jialei, et al.
Pubblicazione: (2025)
Unified Noise Steering for Efficient Human-Guided VLA Adaptation
di: Lu, Junjie, et al.
Pubblicazione: (2026)
di: Lu, Junjie, et al.
Pubblicazione: (2026)
VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts
di: Jiang, Yuhua, et al.
Pubblicazione: (2026)
di: Jiang, Yuhua, et al.
Pubblicazione: (2026)
Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation
di: Zhu, Yihang, et al.
Pubblicazione: (2025)
di: Zhu, Yihang, et al.
Pubblicazione: (2025)
VLA Model Post-Training via Action-Chunked PPO and Self Behavior Cloning
di: Wang, Si-Cheng, et al.
Pubblicazione: (2025)
di: Wang, Si-Cheng, et al.
Pubblicazione: (2025)
Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds
di: Fan, Xianzhe, et al.
Pubblicazione: (2026)
di: Fan, Xianzhe, et al.
Pubblicazione: (2026)
SELF-VLA: A Skill Enhanced Agentic Vision-Language-Action Framework for Contact-Rich Disassembly
di: Liu, Chang, et al.
Pubblicazione: (2026)
di: Liu, Chang, et al.
Pubblicazione: (2026)
Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World
di: Jian, Yingzhao, et al.
Pubblicazione: (2025)
di: Jian, Yingzhao, et al.
Pubblicazione: (2025)
Hydra-NeXt: Robust Closed-Loop Driving with Open-Loop Training
di: Li, Zhenxin, et al.
Pubblicazione: (2025)
di: Li, Zhenxin, et al.
Pubblicazione: (2025)
WorldVLA: Towards Autoregressive Action World Model
di: Cen, Jun, et al.
Pubblicazione: (2025)
di: Cen, Jun, et al.
Pubblicazione: (2025)
FrameSkip: Learning from Fewer but More Informative Frames in VLA Training
di: Yu, Bin, et al.
Pubblicazione: (2026)
di: Yu, Bin, et al.
Pubblicazione: (2026)
Dexterous Grasp Transformer
di: Xu, Guo-Hao, et al.
Pubblicazione: (2024)
di: Xu, Guo-Hao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
DHAGrasp: Synthesizing Affordance-Aware Dual-Hand Grasps with Text Instructions
di: Li, Quanzhou, et al.
Pubblicazione: (2025) -
MOWA: Multiple-in-One Image Warping Model
di: Liao, Kang, et al.
Pubblicazione: (2024) -
Decomposed Object Manipulation via Dual-Actor Policy
di: Fan, Bin, et al.
Pubblicazione: (2025) -
OHP-RL: Online Human Preference as Guidance in Reinforcement Learning for Robot Manipulation
di: Mo, Yunyang, et al.
Pubblicazione: (2026) -
Real-to-Sim Grasp: Rethinking the Gap between Simulation and Real World in Grasp Detection
di: Cai, Jia-Feng, et al.
Pubblicazione: (2024)