Mixture of Horizons in Action Chunking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jing, Dong, Wang, Gang, Liu, Jiaqi, Tang, Weiliang, Sun, Zelong, Yao, Yunchao, Wei, Zhenyu, Liu, Yunhui, Lu, Zhiwu, Ding, Mingyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
von: Fang, Yu, et al.
Veröffentlicht: (2026)
von: Fang, Yu, et al.
Veröffentlicht: (2026)
REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation
von: Yuan, Puzhen, et al.
Veröffentlicht: (2025)
von: Yuan, Puzhen, et al.
Veröffentlicht: (2025)
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval
von: Sun, Zelong, et al.
Veröffentlicht: (2025)
von: Sun, Zelong, et al.
Veröffentlicht: (2025)
One Hand to Rule Them All: Canonical Representations for Unified Dexterous Manipulation
von: Wei, Zhenyu, et al.
Veröffentlicht: (2026)
von: Wei, Zhenyu, et al.
Veröffentlicht: (2026)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion Policy
von: Wu, Pengyuan, et al.
Veröffentlicht: (2026)
von: Wu, Pengyuan, et al.
Veröffentlicht: (2026)
ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks
von: Wang, Kaijun, et al.
Veröffentlicht: (2025)
von: Wang, Kaijun, et al.
Veröffentlicht: (2025)
Say Cheese! Detail-Preserving Portrait Collection Generation via Natural Language Edits
von: Sun, Zelong, et al.
Veröffentlicht: (2026)
von: Sun, Zelong, et al.
Veröffentlicht: (2026)
BOSS: Benchmark for Observation Space Shift in Long-Horizon Task
von: Yang, Yue, et al.
Veröffentlicht: (2025)
von: Yang, Yue, et al.
Veröffentlicht: (2025)
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
von: Wang, Yating, et al.
Veröffentlicht: (2025)
von: Wang, Yating, et al.
Veröffentlicht: (2025)
SldprtNet: A Large-Scale Multimodal Dataset for CAD Generation in Language-Driven 3D Design
von: Li, Ruogu, et al.
Veröffentlicht: (2026)
von: Li, Ruogu, et al.
Veröffentlicht: (2026)
World Guidance: World Modeling in Condition Space for Action Generation
von: Su, Yue, et al.
Veröffentlicht: (2026)
von: Su, Yue, et al.
Veröffentlicht: (2026)
From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation
von: Li, Yajie, et al.
Veröffentlicht: (2026)
von: Li, Yajie, et al.
Veröffentlicht: (2026)
Unifying Language-Action Understanding and Generation for Autonomous Driving
von: Wang, Xinyang, et al.
Veröffentlicht: (2026)
von: Wang, Xinyang, et al.
Veröffentlicht: (2026)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
von: Chen, Yi, et al.
Veröffentlicht: (2024)
von: Chen, Yi, et al.
Veröffentlicht: (2024)
InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual Manipulation
von: Lee, Andrew, et al.
Veröffentlicht: (2024)
von: Lee, Andrew, et al.
Veröffentlicht: (2024)
Leave No Observation Behind: Real-time Correction for VLA Action Chunks
von: Sendai, Kohei, et al.
Veröffentlicht: (2025)
von: Sendai, Kohei, et al.
Veröffentlicht: (2025)
Enhanced Scale-aware Depth Estimation for Monocular Endoscopic Scenes with Geometric Modeling
von: Wei, Ruofeng, et al.
Veröffentlicht: (2024)
von: Wei, Ruofeng, et al.
Veröffentlicht: (2024)
DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA
von: Chen, Yi, et al.
Veröffentlicht: (2026)
von: Chen, Yi, et al.
Veröffentlicht: (2026)
Efficient Long-Horizon Vision-Language-Action Models via Static-Dynamic Disentanglement
von: Qiu, Weikang, et al.
Veröffentlicht: (2026)
von: Qiu, Weikang, et al.
Veröffentlicht: (2026)
Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition
von: Wen, Yuhang, et al.
Veröffentlicht: (2023)
von: Wen, Yuhang, et al.
Veröffentlicht: (2023)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
von: Li, Runhao, et al.
Veröffentlicht: (2025)
von: Li, Runhao, et al.
Veröffentlicht: (2025)
General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling
von: Lyu, Huaihai, et al.
Veröffentlicht: (2026)
von: Lyu, Huaihai, et al.
Veröffentlicht: (2026)
Bridge Thinking and Acting: Unleashing Physical Potential of VLM with Generalizable Action Expert
von: Liu, Mingyu, et al.
Veröffentlicht: (2025)
von: Liu, Mingyu, et al.
Veröffentlicht: (2025)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
von: Liang, Wenqi, et al.
Veröffentlicht: (2025)
von: Liang, Wenqi, et al.
Veröffentlicht: (2025)
Closed Loop Dynamic Driving Data Mixture for Real-Synthetic Co-Training
von: Ruan, Hongzhi, et al.
Veröffentlicht: (2026)
von: Ruan, Hongzhi, et al.
Veröffentlicht: (2026)
LongNav-R1: Horizon-Adaptive Multi-Turn RL for Long-Horizon VLA Navigation
von: Hu, Yue, et al.
Veröffentlicht: (2026)
von: Hu, Yue, et al.
Veröffentlicht: (2026)
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
von: Zhao, Chen, et al.
Veröffentlicht: (2026)
von: Zhao, Chen, et al.
Veröffentlicht: (2026)
GraSP-VLA: Graph-based Symbolic Action Representation for Long-Horizon Planning with VLA Policies
von: Neau, Maëlic, et al.
Veröffentlicht: (2025)
von: Neau, Maëlic, et al.
Veröffentlicht: (2025)
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
von: Tang, Zuojin, et al.
Veröffentlicht: (2026)
von: Tang, Zuojin, et al.
Veröffentlicht: (2026)
StructBiHOI: Structured Articulation Modeling for Long--Horizon Bimanual Hand--Object Interaction Generation
von: Wang, Zhi, et al.
Veröffentlicht: (2026)
von: Wang, Zhi, et al.
Veröffentlicht: (2026)
Action Emergence from Streaming Intent
von: Jing, Pengfei, et al.
Veröffentlicht: (2026)
von: Jing, Pengfei, et al.
Veröffentlicht: (2026)
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
von: Huang, Zheng, et al.
Veröffentlicht: (2025)
von: Huang, Zheng, et al.
Veröffentlicht: (2025)
SCAR: Self-Supervised Continuous Action Representation Learning
von: Liu, Hongjia, et al.
Veröffentlicht: (2026)
von: Liu, Hongjia, et al.
Veröffentlicht: (2026)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
von: Liu, Yicheng, et al.
Veröffentlicht: (2025)
von: Liu, Yicheng, et al.
Veröffentlicht: (2025)
Adapting Pre-Trained Vision Models for Novel Instance Detection and Segmentation
von: Lu, Yangxiao, et al.
Veröffentlicht: (2024)
von: Lu, Yangxiao, et al.
Veröffentlicht: (2024)
Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
von: Liu, Dayong, et al.
Veröffentlicht: (2025)
von: Liu, Dayong, et al.
Veröffentlicht: (2025)
GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
von: Fang, Yu, et al.
Veröffentlicht: (2026) -
REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation
von: Yuan, Puzhen, et al.
Veröffentlicht: (2025) -
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval
von: Sun, Zelong, et al.
Veröffentlicht: (2025) -
One Hand to Rule Them All: Canonical Representations for Unified Dexterous Manipulation
von: Wei, Zhenyu, et al.
Veröffentlicht: (2026) -
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)