HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Minghui, Ding, Pengxiang, Wang, Shu, Zhuang, Zifeng, Liu, Yang, Tong, Xinyang, Song, Wenxuan, Lyu, Shangke, Huang, Siteng, Wang, Donglin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
by: Ding, Pengxiang, et al.
Published: (2023)
by: Ding, Pengxiang, et al.
Published: (2023)
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
by: Zhao, Han, et al.
Published: (2025)
by: Zhao, Han, et al.
Published: (2025)
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
by: Wang, Yihao, et al.
Published: (2025)
by: Wang, Yihao, et al.
Published: (2025)
World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
by: Li, Runze, et al.
Published: (2026)
by: Li, Runze, et al.
Published: (2026)
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models
by: Zhao, Han, et al.
Published: (2025)
by: Zhao, Han, et al.
Published: (2025)
CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
by: Fan, Yiguo, et al.
Published: (2025)
by: Fan, Yiguo, et al.
Published: (2025)
HiF-DTA: Hierarchical Feature Learning Network for Drug-Target Affinity Prediction
by: Li, Minghui, et al.
Published: (2025)
by: Li, Minghui, et al.
Published: (2025)
GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation
by: Zhang, Hongyin, et al.
Published: (2025)
by: Zhang, Hongyin, et al.
Published: (2025)
Robust Online Residual Refinement via Koopman-Guided Dynamics Modeling
by: Gong, Zhefei, et al.
Published: (2025)
by: Gong, Zhefei, et al.
Published: (2025)
CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
by: Gong, Zhefei, et al.
Published: (2024)
by: Gong, Zhefei, et al.
Published: (2024)
GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot
by: Song, Wenxuan, et al.
Published: (2024)
by: Song, Wenxuan, et al.
Published: (2024)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
by: Chen, Jiayi, et al.
Published: (2025)
by: Chen, Jiayi, et al.
Published: (2025)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only
by: Xiao, Wei, et al.
Published: (2025)
by: Xiao, Wei, et al.
Published: (2025)
QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning
by: Tong, Xinyang, et al.
Published: (2024)
by: Tong, Xinyang, et al.
Published: (2024)
ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
by: Cui, Can, et al.
Published: (2024)
by: Cui, Can, et al.
Published: (2024)
OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
by: Cui, Can, et al.
Published: (2025)
by: Cui, Can, et al.
Published: (2025)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
Dynamic Adaptive Legged Locomotion Policy via Decoupling Reaction Force Control and Gait Control
by: Wang, Renjie, et al.
Published: (2025)
by: Wang, Renjie, et al.
Published: (2025)
Discovering Self-Protective Falling Policy for Humanoid Robot via Deep Reinforcement Learning
by: Shi, Diyuan, et al.
Published: (2025)
by: Shi, Diyuan, et al.
Published: (2025)
HiFo-Prompt: Prompting with Hindsight and Foresight for LLM-based Automatic Heuristic Design
by: Chen, Chentong, et al.
Published: (2025)
by: Chen, Chentong, et al.
Published: (2025)
RationalVLA: A Rational Vision-Language-Action Model with Dual System
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
TDMPBC: Self-Imitative Reinforcement Learning for Humanoid Robot Control
by: Zhuang, Zifeng, et al.
Published: (2025)
by: Zhuang, Zifeng, et al.
Published: (2025)
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
by: Song, Wenxuan, et al.
Published: (2026)
by: Song, Wenxuan, et al.
Published: (2026)
Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration
by: Ding, Pengxiang, et al.
Published: (2025)
by: Ding, Pengxiang, et al.
Published: (2025)
ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning
by: Zhang, Hongyin, et al.
Published: (2025)
by: Zhang, Hongyin, et al.
Published: (2025)
VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
by: Li, Hengtao, et al.
Published: (2025)
by: Li, Hengtao, et al.
Published: (2025)
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
by: Li, Fuhao, et al.
Published: (2025)
by: Li, Fuhao, et al.
Published: (2025)
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
by: Song, Wenxuan, et al.
Published: (2026)
by: Song, Wenxuan, et al.
Published: (2026)
Activation-wise Propagation: A One-Timestep Strategy for Spiking Neural Networks
by: Song, Jian, et al.
Published: (2025)
by: Song, Jian, et al.
Published: (2025)
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference
by: Zhao, Han, et al.
Published: (2024)
by: Zhao, Han, et al.
Published: (2024)
SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Exploring the Evolution of Physics Cognition in Video Generation: A Survey
by: Lin, Minghui, et al.
Published: (2025)
by: Lin, Minghui, et al.
Published: (2025)
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
by: Zhao, Wei, et al.
Published: (2025)
by: Zhao, Wei, et al.
Published: (2025)
Unlock Reliable Skill Inference for Quadruped Adaptive Behavior by Skill Graph
by: Zhang, Hongyin, et al.
Published: (2023)
by: Zhang, Hongyin, et al.
Published: (2023)
VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL
by: Dai, Fengyuan, et al.
Published: (2025)
by: Dai, Fengyuan, et al.
Published: (2025)
From Hindsight to Foresight: Self-Encouraged Hindsight Distillation for Knowledge-based Visual Question Answering
by: Zhao, Yu, et al.
Published: (2025)
by: Zhao, Yu, et al.
Published: (2025)
Similar Items
-
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
by: Ding, Pengxiang, et al.
Published: (2023) -
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
by: Liu, Yang, et al.
Published: (2026) -
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
by: Zhao, Han, et al.
Published: (2025) -
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
by: Wang, Yihao, et al.
Published: (2025) -
World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
by: Li, Runze, et al.
Published: (2026)