BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Peiyan, Chen, Yixiang, Wu, Hongtao, Ma, Xiao, Wu, Xiangnan, Huang, Yan, Wang, Liang, Kong, Tao, Tan, Tieniu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation
di: Chen, Yixiang, et al.
Pubblicazione: (2025)
di: Chen, Yixiang, et al.
Pubblicazione: (2025)
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model
di: Li, Peiyan, et al.
Pubblicazione: (2026)
di: Li, Peiyan, et al.
Pubblicazione: (2026)
GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy
di: Li, Peiyan, et al.
Pubblicazione: (2024)
di: Li, Peiyan, et al.
Pubblicazione: (2024)
BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks
di: Chen, Yixiang, et al.
Pubblicazione: (2026)
di: Chen, Yixiang, et al.
Pubblicazione: (2026)
TaF-VLA: Tactile-Force Alignment in Vision-Language-Action Models for Force-aware Manipulation
di: Huang, Yuzhe, et al.
Pubblicazione: (2026)
di: Huang, Yuzhe, et al.
Pubblicazione: (2026)
EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow
di: Chen, Yixiang, et al.
Pubblicazione: (2025)
di: Chen, Yixiang, et al.
Pubblicazione: (2025)
UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models
di: Yang, Jiabing, et al.
Pubblicazione: (2026)
di: Yang, Jiabing, et al.
Pubblicazione: (2026)
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
di: Peng, Xiongfeng, et al.
Pubblicazione: (2026)
di: Peng, Xiongfeng, et al.
Pubblicazione: (2026)
AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment
di: Kong, Weijie, et al.
Pubblicazione: (2026)
di: Kong, Weijie, et al.
Pubblicazione: (2026)
FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation
di: Miao, Cui, et al.
Pubblicazione: (2025)
di: Miao, Cui, et al.
Pubblicazione: (2025)
ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models
di: Li, Ye, et al.
Pubblicazione: (2026)
di: Li, Ye, et al.
Pubblicazione: (2026)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
di: Li, Runhao, et al.
Pubblicazione: (2025)
di: Li, Runhao, et al.
Pubblicazione: (2025)
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
di: Sun, Jianli, et al.
Pubblicazione: (2026)
di: Sun, Jianli, et al.
Pubblicazione: (2026)
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
di: Xu, Siyu, et al.
Pubblicazione: (2025)
di: Xu, Siyu, et al.
Pubblicazione: (2025)
ProgressVLA: Progress-Guided Diffusion Policy for Vision-Language Robotic Manipulation
di: Yan, Hongyu, et al.
Pubblicazione: (2026)
di: Yan, Hongyu, et al.
Pubblicazione: (2026)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
di: Wen, Junjie, et al.
Pubblicazione: (2024)
di: Wen, Junjie, et al.
Pubblicazione: (2024)
GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
di: Huang, Helong, et al.
Pubblicazione: (2025)
di: Huang, Helong, et al.
Pubblicazione: (2025)
EchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation
di: Lin, Min, et al.
Pubblicazione: (2025)
di: Lin, Min, et al.
Pubblicazione: (2025)
MoManipVLA: Transferring Vision-language-action Models for General Mobile Manipulation
di: Wu, Zhenyu, et al.
Pubblicazione: (2025)
di: Wu, Zhenyu, et al.
Pubblicazione: (2025)
SKIP: Sparse Keyframe Interpolation Paradigm for Efficient Embodied World Models
di: He, Ziheng, et al.
Pubblicazione: (2026)
di: He, Ziheng, et al.
Pubblicazione: (2026)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
di: Li, Qiwei, et al.
Pubblicazione: (2026)
di: Li, Qiwei, et al.
Pubblicazione: (2026)
ST-$π$: Structured SpatioTemporal VLA for Robotic Manipulation
di: Ma, Chuanhao, et al.
Pubblicazione: (2026)
di: Ma, Chuanhao, et al.
Pubblicazione: (2026)
EgoDemoGen: Egocentric Demonstration Generation for Viewpoint Generalization in Robotic Manipulation
di: Xu, Yuan, et al.
Pubblicazione: (2025)
di: Xu, Yuan, et al.
Pubblicazione: (2025)
IA-VLA: Input Augmentation for Vision-Language-Action models in settings with semantically complex tasks
di: Hannus, Eric, et al.
Pubblicazione: (2025)
di: Hannus, Eric, et al.
Pubblicazione: (2025)
$M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills
di: Xiao, Siyao, et al.
Pubblicazione: (2026)
di: Xiao, Siyao, et al.
Pubblicazione: (2026)
OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
di: Liu, Ruixun, et al.
Pubblicazione: (2025)
di: Liu, Ruixun, et al.
Pubblicazione: (2025)
AnchorVLA4D: an Anchor-Based Spatial-Temporal Vision-Language-Action Model for Robotic Manipulation
di: Zhu, Juan, et al.
Pubblicazione: (2026)
di: Zhu, Juan, et al.
Pubblicazione: (2026)
ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
di: Liu, Zhenyang, et al.
Pubblicazione: (2026)
di: Liu, Zhenyang, et al.
Pubblicazione: (2026)
TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
di: Im, Hokyun, et al.
Pubblicazione: (2025)
di: Im, Hokyun, et al.
Pubblicazione: (2025)
SG-VLA: Learning Spatially-Grounded Vision-Language-Action Models for Mobile Manipulation
di: Tu, Ruisen, et al.
Pubblicazione: (2026)
di: Tu, Ruisen, et al.
Pubblicazione: (2026)
TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation
di: Zhang, Kaidi, et al.
Pubblicazione: (2026)
di: Zhang, Kaidi, et al.
Pubblicazione: (2026)
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
di: Zhang, Rongyu, et al.
Pubblicazione: (2025)
di: Zhang, Rongyu, et al.
Pubblicazione: (2025)
VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation
di: Zhong, Zhide, et al.
Pubblicazione: (2026)
di: Zhong, Zhide, et al.
Pubblicazione: (2026)
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
di: Hu, Xintong, et al.
Pubblicazione: (2026)
di: Hu, Xintong, et al.
Pubblicazione: (2026)
DySL-VLA: Efficient Vision-Language-Action Model Inference via Dynamic-Static Layer-Skipping for Robot Manipulation
di: Yang, Zebin, et al.
Pubblicazione: (2026)
di: Yang, Zebin, et al.
Pubblicazione: (2026)
EdgeVLA: Efficient Vision-Language-Action Models
di: Budzianowski, Paweł, et al.
Pubblicazione: (2025)
di: Budzianowski, Paweł, et al.
Pubblicazione: (2025)
SimVLA: A Simple VLA Baseline for Robotic Manipulation
di: Luo, Yuankai, et al.
Pubblicazione: (2026)
di: Luo, Yuankai, et al.
Pubblicazione: (2026)
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
di: Yu, Wenda, et al.
Pubblicazione: (2026)
di: Yu, Wenda, et al.
Pubblicazione: (2026)
SELF-VLA: A Skill Enhanced Agentic Vision-Language-Action Framework for Contact-Rich Disassembly
di: Liu, Chang, et al.
Pubblicazione: (2026)
di: Liu, Chang, et al.
Pubblicazione: (2026)
DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
di: Su, Taiyi, et al.
Pubblicazione: (2026)
di: Su, Taiyi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation
di: Chen, Yixiang, et al.
Pubblicazione: (2025) -
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model
di: Li, Peiyan, et al.
Pubblicazione: (2026) -
GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy
di: Li, Peiyan, et al.
Pubblicazione: (2024) -
BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks
di: Chen, Yixiang, et al.
Pubblicazione: (2026) -
TaF-VLA: Tactile-Force Alignment in Vision-Language-Action Models for Force-aware Manipulation
di: Huang, Yuzhe, et al.
Pubblicazione: (2026)