Beyond World-Frame Action Heads: Motion-Centric Action Frames for Vision-Language-Action Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Huoren, Zhao, Jianchao, Yusong, Hu, Ou, Qiguan, Gao, Yuyang, Ke, Wei, He, Yuhang, Dong, SongLin, Ma, Zhiheng, Gong, Yihong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs
di: Zhao, Jianchao, et al.
Pubblicazione: (2026)
di: Zhao, Jianchao, et al.
Pubblicazione: (2026)
Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models
di: Liu, Haoyun, et al.
Pubblicazione: (2026)
di: Liu, Haoyun, et al.
Pubblicazione: (2026)
ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning
di: Nazarenus, Eric, et al.
Pubblicazione: (2026)
di: Nazarenus, Eric, et al.
Pubblicazione: (2026)
Unleashing the Potential of All Test Samples: Mean-Shift Guided Test-Time Adaptation
di: Han, Jizhou, et al.
Pubblicazione: (2025)
di: Han, Jizhou, et al.
Pubblicazione: (2025)
Unify Robot Actions in Camera Frame
di: Xie, Sicheng, et al.
Pubblicazione: (2025)
di: Xie, Sicheng, et al.
Pubblicazione: (2025)
Continuous Expert Assembly: Instance-Conditioned Low-Rank Residuals for All-in-One Image Restoration
di: He, Haisen, et al.
Pubblicazione: (2026)
di: He, Haisen, et al.
Pubblicazione: (2026)
Trajectory-Diversity-Driven Robust Vision-and-Language Navigation
di: Li, Jiangyang, et al.
Pubblicazione: (2026)
di: Li, Jiangyang, et al.
Pubblicazione: (2026)
ReMoT: Reinforcement Learning with Motion Contrast Triplets
di: Wan, Cong, et al.
Pubblicazione: (2026)
di: Wan, Cong, et al.
Pubblicazione: (2026)
Recurrence-Complete Frame-based Action Models
di: Keiblinger, Michael
Pubblicazione: (2025)
di: Keiblinger, Michael
Pubblicazione: (2025)
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
di: Wang, Qiang, et al.
Pubblicazione: (2025)
di: Wang, Qiang, et al.
Pubblicazione: (2025)
ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
di: Jang, Huiwon, et al.
Pubblicazione: (2025)
di: Jang, Huiwon, et al.
Pubblicazione: (2025)
DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions
di: Li, Zongyue, et al.
Pubblicazione: (2025)
di: Li, Zongyue, et al.
Pubblicazione: (2025)
Understanding Multimodal Complementarity for Single-Frame Action Anticipation
di: Benavent-Lledo, Manuel, et al.
Pubblicazione: (2026)
di: Benavent-Lledo, Manuel, et al.
Pubblicazione: (2026)
ReSpike: Residual Frames-based Hybrid Spiking Neural Networks for Efficient Action Recognition
di: Xiao, Shiting, et al.
Pubblicazione: (2024)
di: Xiao, Shiting, et al.
Pubblicazione: (2024)
EventCrab: Harnessing Frame and Point Synergy for Event-based Action Recognition and Beyond
di: Cao, Meiqi, et al.
Pubblicazione: (2024)
di: Cao, Meiqi, et al.
Pubblicazione: (2024)
DualCP: Rehearsal-Free Domain-Incremental Learning via Dual-Level Concept Prototype
di: Wang, Qiang, et al.
Pubblicazione: (2025)
di: Wang, Qiang, et al.
Pubblicazione: (2025)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
di: Zhang, Tianyi, et al.
Pubblicazione: (2025)
di: Zhang, Tianyi, et al.
Pubblicazione: (2025)
GOAL: Geometrically Optimal Alignment for Continual Generalized Category Discovery
di: Han, Jizhou, et al.
Pubblicazione: (2026)
di: Han, Jizhou, et al.
Pubblicazione: (2026)
Learning Like Humans: Analogical Concept Learning for Generalized Category Discovery
di: Han, Jizhou, et al.
Pubblicazione: (2026)
di: Han, Jizhou, et al.
Pubblicazione: (2026)
Consistent Supervised-Unsupervised Alignment for Generalized Category Discovery
di: Han, Jizhou, et al.
Pubblicazione: (2025)
di: Han, Jizhou, et al.
Pubblicazione: (2025)
Permutation-Aware Action Segmentation via Unsupervised Frame-to-Segment Alignment
di: Tran, Quoc-Huy, et al.
Pubblicazione: (2023)
di: Tran, Quoc-Huy, et al.
Pubblicazione: (2023)
Training-Free Action Recognition and Goal Inference with Dynamic Frame Selection
di: Keat, Ee Yeo, et al.
Pubblicazione: (2024)
di: Keat, Ee Yeo, et al.
Pubblicazione: (2024)
One-Frame Calibration with Siamese Network in Facial Action Unit Recognition
di: Feng, Shuangquan, et al.
Pubblicazione: (2024)
di: Feng, Shuangquan, et al.
Pubblicazione: (2024)
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
di: Zhao, Baining, et al.
Pubblicazione: (2026)
di: Zhao, Baining, et al.
Pubblicazione: (2026)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
di: Li, Hao, et al.
Pubblicazione: (2025)
di: Li, Hao, et al.
Pubblicazione: (2025)
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
di: Li, Boyu, et al.
Pubblicazione: (2026)
di: Li, Boyu, et al.
Pubblicazione: (2026)
EgoAction: Egocentric Action Composition with Reliability-Aware Temporal Fusion for the EPIC-KITCHENS Action Detection Challenge at CVPR 2026
di: Fu, Zhiheng, et al.
Pubblicazione: (2026)
di: Fu, Zhiheng, et al.
Pubblicazione: (2026)
Stochastic Human Motion Prediction with Memory of Action Transition and Action Characteristic
di: Tang, Jianwei, et al.
Pubblicazione: (2025)
di: Tang, Jianwei, et al.
Pubblicazione: (2025)
From Frames to Sequences: Temporally Consistent Human-Centric Dense Prediction
di: Miao, Xingyu, et al.
Pubblicazione: (2026)
di: Miao, Xingyu, et al.
Pubblicazione: (2026)
World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
di: Li, Runze, et al.
Pubblicazione: (2026)
di: Li, Runze, et al.
Pubblicazione: (2026)
CASR: Refining Action Segmentation via Marginalizing Frame-levle Causal Relationships
di: Du, Keqing, et al.
Pubblicazione: (2023)
di: Du, Keqing, et al.
Pubblicazione: (2023)
LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning
di: Lai, Bolin, et al.
Pubblicazione: (2023)
di: Lai, Bolin, et al.
Pubblicazione: (2023)
Common Ground: Framing and the Potential to Mitigate Herbicide Resistance Using Collective Action
di: Ariel Singerman, et al.
Pubblicazione: (2025)
di: Ariel Singerman, et al.
Pubblicazione: (2025)
E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving
di: Tang, Yihong, et al.
Pubblicazione: (2025)
di: Tang, Yihong, et al.
Pubblicazione: (2025)
Beyond Prompt Learning: Continual Adapter for Efficient Rehearsal-Free Continual Learning
di: Gao, Xinyuan, et al.
Pubblicazione: (2024)
di: Gao, Xinyuan, et al.
Pubblicazione: (2024)
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
di: Tang, Zuojin, et al.
Pubblicazione: (2026)
di: Tang, Zuojin, et al.
Pubblicazione: (2026)
Vision-Language Models Unlock Task-Centric Latent Actions
di: Nikulin, Alexander, et al.
Pubblicazione: (2026)
di: Nikulin, Alexander, et al.
Pubblicazione: (2026)
Full‐Scale Tests to Characterize the Effect of Framing Action and Slab Continuity on the Collapse Capacity of Composite Frames Under Cyclic Loading
di: Hammad El Jisr, et al.
Pubblicazione: (2024)
di: Hammad El Jisr, et al.
Pubblicazione: (2024)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
di: Li, Qiwei, et al.
Pubblicazione: (2026)
di: Li, Qiwei, et al.
Pubblicazione: (2026)
Light Future: Multimodal Action Frame Prediction via InstructPix2Pix
di: Zhong, Zesen, et al.
Pubblicazione: (2025)
di: Zhong, Zesen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs
di: Zhao, Jianchao, et al.
Pubblicazione: (2026) -
Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models
di: Liu, Haoyun, et al.
Pubblicazione: (2026) -
ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning
di: Nazarenus, Eric, et al.
Pubblicazione: (2026) -
Unleashing the Potential of All Test Samples: Mean-Shift Guided Test-Time Adaptation
di: Han, Jizhou, et al.
Pubblicazione: (2025) -
Unify Robot Actions in Camera Frame
di: Xie, Sicheng, et al.
Pubblicazione: (2025)