Bridging Language, Vision and Action: Multimodal VAEs in Robotic Manipulation Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Sejnova, Gabriela, Vavrecka, Michal, Stepanova, Karla |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PRAG: Procedural Action Generator
di: Vavrecka, Michal, et al.
Pubblicazione: (2025)
di: Vavrecka, Michal, et al.
Pubblicazione: (2025)
Adaptive Compression of the Latent Space in Variational Autoencoders
di: Sejnova, Gabriela, et al.
Pubblicazione: (2023)
di: Sejnova, Gabriela, et al.
Pubblicazione: (2023)
Benchmarking Multimodal Variational Autoencoders: CdSprites+ Dataset and Toolkit
di: Sejnova, Gabriela, et al.
Pubblicazione: (2022)
di: Sejnova, Gabriela, et al.
Pubblicazione: (2022)
TransforMerger: Transformer-based Voice-Gesture Fusion for Robust Human-Robot Communication
di: Vanc, Petr, et al.
Pubblicazione: (2025)
di: Vanc, Petr, et al.
Pubblicazione: (2025)
Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots
di: Su, Haochen, et al.
Pubblicazione: (2025)
di: Su, Haochen, et al.
Pubblicazione: (2025)
Semantic-Geometric Task Representations for Bimanual Manipulation from Human Demonstrations to Robot Action Planning
di: Herbert, Franziska, et al.
Pubblicazione: (2026)
di: Herbert, Franziska, et al.
Pubblicazione: (2026)
Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks
di: Guruprasad, Pranav, et al.
Pubblicazione: (2024)
di: Guruprasad, Pranav, et al.
Pubblicazione: (2024)
Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
di: Guan, Weifan, et al.
Pubblicazione: (2025)
di: Guan, Weifan, et al.
Pubblicazione: (2025)
Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models
di: Francis-Meretzki, Shelly, et al.
Pubblicazione: (2026)
di: Francis-Meretzki, Shelly, et al.
Pubblicazione: (2026)
Language-Conditioned Representations and Mixture-of-Experts Policy for Robust Multi-Task Robotic Manipulation
di: Zhang, Xiucheng, et al.
Pubblicazione: (2025)
di: Zhang, Xiucheng, et al.
Pubblicazione: (2025)
Dynamic Hand Gesture Recognition for Robot Manipulator Tasks
di: Sharma, Dharmendra, et al.
Pubblicazione: (2026)
di: Sharma, Dharmendra, et al.
Pubblicazione: (2026)
OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation
di: Hirose, Noriaki, et al.
Pubblicazione: (2025)
di: Hirose, Noriaki, et al.
Pubblicazione: (2025)
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
di: Shukor, Mustafa, et al.
Pubblicazione: (2025)
di: Shukor, Mustafa, et al.
Pubblicazione: (2025)
Time-Unified Diffusion Policy with Action Discrimination for Robotic Manipulation
di: Niu, Ye, et al.
Pubblicazione: (2025)
di: Niu, Ye, et al.
Pubblicazione: (2025)
MuBlE: MuJoCo and Blender simulation Environment and Benchmark for Task Planning in Robot Manipulation
di: Nazarczuk, Michal, et al.
Pubblicazione: (2025)
di: Nazarczuk, Michal, et al.
Pubblicazione: (2025)
TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
di: Im, Hokyun, et al.
Pubblicazione: (2025)
di: Im, Hokyun, et al.
Pubblicazione: (2025)
DexCanvas: Bridging Human Demonstrations and Robot Learning for Dexterous Manipulation
di: Xu, Xinyue, et al.
Pubblicazione: (2025)
di: Xu, Xinyue, et al.
Pubblicazione: (2025)
$π_0$: A Vision-Language-Action Flow Model for General Robot Control
di: Black, Kevin, et al.
Pubblicazione: (2024)
di: Black, Kevin, et al.
Pubblicazione: (2024)
On the Role of the Action Space in Robot Manipulation Learning and Sim-to-Real Transfer
di: Aljalbout, Elie, et al.
Pubblicazione: (2023)
di: Aljalbout, Elie, et al.
Pubblicazione: (2023)
Interpretable Robotic Manipulation from Language
di: Zheng, Boyuan, et al.
Pubblicazione: (2024)
di: Zheng, Boyuan, et al.
Pubblicazione: (2024)
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025)
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025)
Autoregressive Action Sequence Learning for Robotic Manipulation
di: Zhang, Xinyu, et al.
Pubblicazione: (2024)
di: Zhang, Xinyu, et al.
Pubblicazione: (2024)
RoboBERT: An End-to-end Multimodal Robotic Manipulation Model
di: Wang, Sicheng, et al.
Pubblicazione: (2025)
di: Wang, Sicheng, et al.
Pubblicazione: (2025)
Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
di: Cai, Rui, et al.
Pubblicazione: (2026)
di: Cai, Rui, et al.
Pubblicazione: (2026)
FAST: Efficient Action Tokenization for Vision-Language-Action Models
di: Pertsch, Karl, et al.
Pubblicazione: (2025)
di: Pertsch, Karl, et al.
Pubblicazione: (2025)
Intrinsic Language-Guided Exploration for Complex Long-Horizon Robotic Manipulation Tasks
di: Triantafyllidis, Eleftherios, et al.
Pubblicazione: (2023)
di: Triantafyllidis, Eleftherios, et al.
Pubblicazione: (2023)
DINOBot: Robot Manipulation via Retrieval and Alignment with Vision Foundation Models
di: Di Palo, Norman, et al.
Pubblicazione: (2024)
di: Di Palo, Norman, et al.
Pubblicazione: (2024)
Enhancing Robotic Manipulation with AI Feedback from Multimodal Large Language Models
di: Liu, Jinyi, et al.
Pubblicazione: (2024)
di: Liu, Jinyi, et al.
Pubblicazione: (2024)
Per-Group Error, Not Total MSE: Fine-Tuning Vision-Language-Action Models for 11-DoF Mobile Manipulation
di: Bofi, Pau Montagut, et al.
Pubblicazione: (2026)
di: Bofi, Pau Montagut, et al.
Pubblicazione: (2026)
Confidence Calibration in Vision-Language-Action Models
di: Zollo, Thomas P, et al.
Pubblicazione: (2025)
di: Zollo, Thomas P, et al.
Pubblicazione: (2025)
From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation
di: Yuan, Yifu, et al.
Pubblicazione: (2025)
di: Yuan, Yifu, et al.
Pubblicazione: (2025)
VLP: Vision-Language Preference Learning for Embodied Manipulation
di: Liu, Runze, et al.
Pubblicazione: (2025)
di: Liu, Runze, et al.
Pubblicazione: (2025)
DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization
di: Lin, Sixu, et al.
Pubblicazione: (2026)
di: Lin, Sixu, et al.
Pubblicazione: (2026)
World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
di: Li, Runze, et al.
Pubblicazione: (2026)
di: Li, Runze, et al.
Pubblicazione: (2026)
Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation
di: Feng, Yunhai, et al.
Pubblicazione: (2025)
di: Feng, Yunhai, et al.
Pubblicazione: (2025)
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
di: Li, Qixiu, et al.
Pubblicazione: (2024)
di: Li, Qixiu, et al.
Pubblicazione: (2024)
Act-Observe-Rewrite: Multimodal Coding Agents as In-Context Policy Learners for Robot Manipulation
di: Kumar, Vaishak
Pubblicazione: (2026)
di: Kumar, Vaishak
Pubblicazione: (2026)
Toward Accurate Long-Horizon Robotic Manipulation: Language-to-Action with Foundation Models via Scene Graphs
di: Dinesh, Sushil Samuel, et al.
Pubblicazione: (2025)
di: Dinesh, Sushil Samuel, et al.
Pubblicazione: (2025)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
di: Li, Qixiu, et al.
Pubblicazione: (2025)
di: Li, Qixiu, et al.
Pubblicazione: (2025)
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
di: Shi, Lucy Xiaoyang, et al.
Pubblicazione: (2025)
di: Shi, Lucy Xiaoyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
PRAG: Procedural Action Generator
di: Vavrecka, Michal, et al.
Pubblicazione: (2025) -
Adaptive Compression of the Latent Space in Variational Autoencoders
di: Sejnova, Gabriela, et al.
Pubblicazione: (2023) -
Benchmarking Multimodal Variational Autoencoders: CdSprites+ Dataset and Toolkit
di: Sejnova, Gabriela, et al.
Pubblicazione: (2022) -
TransforMerger: Transformer-based Voice-Gesture Fusion for Robust Human-Robot Communication
di: Vanc, Petr, et al.
Pubblicazione: (2025) -
Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots
di: Su, Haochen, et al.
Pubblicazione: (2025)