GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot
Fuente:
arXiv
Salvato in:
| Autori principali: | Song, Wenxuan, Zhao, Han, Ding, Pengxiang, Cui, Can, Lyu, Shangke, Fan, Yaning, Wang, Donglin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning
di: Tong, Xinyang, et al.
Pubblicazione: (2024)
di: Tong, Xinyang, et al.
Pubblicazione: (2024)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
di: Ding, Pengxiang, et al.
Pubblicazione: (2023)
di: Ding, Pengxiang, et al.
Pubblicazione: (2023)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
di: Gong, Zhefei, et al.
Pubblicazione: (2024)
di: Gong, Zhefei, et al.
Pubblicazione: (2024)
OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
di: Cui, Can, et al.
Pubblicazione: (2025)
di: Cui, Can, et al.
Pubblicazione: (2025)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
di: Chen, Jiayi, et al.
Pubblicazione: (2025)
di: Chen, Jiayi, et al.
Pubblicazione: (2025)
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
di: Song, Wenxuan, et al.
Pubblicazione: (2026)
di: Song, Wenxuan, et al.
Pubblicazione: (2026)
GeRM: A Generative Rendering Model From Physically Realistic to Photorealistic
di: Lu, Jiayuan, et al.
Pubblicazione: (2026)
di: Lu, Jiayuan, et al.
Pubblicazione: (2026)
Discovering Self-Protective Falling Policy for Humanoid Robot via Deep Reinforcement Learning
di: Shi, Diyuan, et al.
Pubblicazione: (2025)
di: Shi, Diyuan, et al.
Pubblicazione: (2025)
Robust Online Residual Refinement via Koopman-Guided Dynamics Modeling
di: Gong, Zhefei, et al.
Pubblicazione: (2025)
di: Gong, Zhefei, et al.
Pubblicazione: (2025)
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models
di: Zhao, Han, et al.
Pubblicazione: (2025)
di: Zhao, Han, et al.
Pubblicazione: (2025)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
The One RING: a Robotic Indoor Navigation Generalist
di: Eftekhar, Ainaz, et al.
Pubblicazione: (2024)
di: Eftekhar, Ainaz, et al.
Pubblicazione: (2024)
GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation
di: Zhang, Hongyin, et al.
Pubblicazione: (2025)
di: Zhang, Hongyin, et al.
Pubblicazione: (2025)
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
di: Song, Wenxuan, et al.
Pubblicazione: (2026)
di: Song, Wenxuan, et al.
Pubblicazione: (2026)
Turning Video Models into Generalist Robot Policies
di: Li, Sizhe Lester, et al.
Pubblicazione: (2026)
di: Li, Sizhe Lester, et al.
Pubblicazione: (2026)
What Matters in Building Vision-Language-Action Models for Generalist Robots
di: Li, Xinghang, et al.
Pubblicazione: (2024)
di: Li, Xinghang, et al.
Pubblicazione: (2024)
Panoramic Multimodal Semantic Occupancy Prediction for Quadruped Robots
di: Zhao, Guoqiang, et al.
Pubblicazione: (2026)
di: Zhao, Guoqiang, et al.
Pubblicazione: (2026)
Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
di: Xing, Youguang, et al.
Pubblicazione: (2025)
di: Xing, Youguang, et al.
Pubblicazione: (2025)
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
di: Zhao, Han, et al.
Pubblicazione: (2025)
di: Zhao, Han, et al.
Pubblicazione: (2025)
ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
di: Cui, Can, et al.
Pubblicazione: (2024)
di: Cui, Can, et al.
Pubblicazione: (2024)
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
di: Hu, Yucheng, et al.
Pubblicazione: (2024)
di: Hu, Yucheng, et al.
Pubblicazione: (2024)
ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation
di: Sun, Yu, et al.
Pubblicazione: (2026)
di: Sun, Yu, et al.
Pubblicazione: (2026)
QuaDreamer: Controllable Panoramic Video Generation for Quadruped Robots
di: Wu, Sheng, et al.
Pubblicazione: (2025)
di: Wu, Sheng, et al.
Pubblicazione: (2025)
Learning Robotic Policy with Imagined Transition: Mitigating the Trade-off between Robustness and Optimality
di: Xiao, Wei, et al.
Pubblicazione: (2025)
di: Xiao, Wei, et al.
Pubblicazione: (2025)
Occupancy World Model for Robots
di: Zhang, Zhang, et al.
Pubblicazione: (2025)
di: Zhang, Zhang, et al.
Pubblicazione: (2025)
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
di: Lin, Minghui, et al.
Pubblicazione: (2025)
di: Lin, Minghui, et al.
Pubblicazione: (2025)
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching
di: Chen, Jiayi, et al.
Pubblicazione: (2026)
di: Chen, Jiayi, et al.
Pubblicazione: (2026)
Unlock Reliable Skill Inference for Quadruped Adaptive Behavior by Skill Graph
di: Zhang, Hongyin, et al.
Pubblicazione: (2023)
di: Zhang, Hongyin, et al.
Pubblicazione: (2023)
ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning
di: Van Vo, Tuan, et al.
Pubblicazione: (2026)
di: Van Vo, Tuan, et al.
Pubblicazione: (2026)
DogSurf: Quadruped Robot Capable of GRU-based Surface Recognition for Blind Person Navigation
di: Bazhenov, Artem, et al.
Pubblicazione: (2024)
di: Bazhenov, Artem, et al.
Pubblicazione: (2024)
Dynamic Adaptive Legged Locomotion Policy via Decoupling Reaction Force Control and Gait Control
di: Wang, Renjie, et al.
Pubblicazione: (2025)
di: Wang, Renjie, et al.
Pubblicazione: (2025)
AgriVLN: Vision-and-Language Navigation for Agricultural Robots
di: Zhao, Xiaobei, et al.
Pubblicazione: (2025)
di: Zhao, Xiaobei, et al.
Pubblicazione: (2025)
DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
di: Gao, Shenyuan, et al.
Pubblicazione: (2026)
di: Gao, Shenyuan, et al.
Pubblicazione: (2026)
RobotPan: A 360$^\circ$ Surround-View Robotic Vision System for Embodied Perception
di: Ma, Jiahao, et al.
Pubblicazione: (2026)
di: Ma, Jiahao, et al.
Pubblicazione: (2026)
VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
di: Zhao, Wei, et al.
Pubblicazione: (2025)
di: Zhao, Wei, et al.
Pubblicazione: (2025)
Activation-wise Propagation: A One-Timestep Strategy for Spiking Neural Networks
di: Song, Jian, et al.
Pubblicazione: (2025)
di: Song, Jian, et al.
Pubblicazione: (2025)
ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning
di: Zhang, Hongyin, et al.
Pubblicazione: (2025)
di: Zhang, Hongyin, et al.
Pubblicazione: (2025)
Radar and Event Camera Fusion for Agile Robot Ego-Motion Estimation
di: Lyu, Yang, et al.
Pubblicazione: (2025)
di: Lyu, Yang, et al.
Pubblicazione: (2025)
From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation
di: Li, Yajie, et al.
Pubblicazione: (2026)
di: Li, Yajie, et al.
Pubblicazione: (2026)
Documenti analoghi
-
QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning
di: Tong, Xinyang, et al.
Pubblicazione: (2024) -
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
di: Ding, Pengxiang, et al.
Pubblicazione: (2023) -
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
di: Song, Wenxuan, et al.
Pubblicazione: (2025) -
CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
di: Gong, Zhefei, et al.
Pubblicazione: (2024) -
OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
di: Cui, Can, et al.
Pubblicazione: (2025)