S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Haodong, Zhong, Zhide, Zhu, Jiaguan, He, Junjie, Yuan, Weilin, Song, Wenxuan, Gong, Xin, Cai, Yingjie, Zhao, Guanyi, Yan, Xu, Liu, Bingbing, Chen, Ying-Cong, Li, Haoang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation
by: Zhong, Zhide, et al.
Published: (2026)
by: Zhong, Zhide, et al.
Published: (2026)
DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models
by: Zhong, Zhide, et al.
Published: (2026)
by: Zhong, Zhide, et al.
Published: (2026)
Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation
by: Yan, Haodong, et al.
Published: (2025)
by: Yan, Haodong, et al.
Published: (2025)
FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
by: Zhong, Zhide, et al.
Published: (2025)
by: Zhong, Zhide, et al.
Published: (2025)
SyntheOcc: Synthesize Geometric-Controlled Street View Images through 3D Semantic MPIs
by: Li, Leheng, et al.
Published: (2024)
by: Li, Leheng, et al.
Published: (2024)
Privileged Foresight Distillation: Zero-Cost Future Correction for World Action Models
by: Fang, Pengcheng, et al.
Published: (2026)
by: Fang, Pengcheng, et al.
Published: (2026)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
Rethinking the Practicality of Vision-language-action Model: A Comprehensive Benchmark and An Improved Baseline
by: Song, Wenxuan, et al.
Published: (2026)
by: Song, Wenxuan, et al.
Published: (2026)
From Hindsight to Foresight: Self-Encouraged Hindsight Distillation for Knowledge-based Visual Question Answering
by: Zhao, Yu, et al.
Published: (2025)
by: Zhao, Yu, et al.
Published: (2025)
GLaD: Geometric Latent Distillation for Vision-Language-Action Models
by: Guo, Minghao, et al.
Published: (2025)
by: Guo, Minghao, et al.
Published: (2025)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
by: Cai, Weilin, et al.
Published: (2024)
by: Cai, Weilin, et al.
Published: (2024)
Lensless and Lossless HoloVAM
by: Madsen, Andreas Erik Gejl, et al.
Published: (2025)
by: Madsen, Andreas Erik Gejl, et al.
Published: (2025)
Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
by: He, Jing, et al.
Published: (2025)
by: He, Jing, et al.
Published: (2025)
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching
by: Chen, Jiayi, et al.
Published: (2026)
by: Chen, Jiayi, et al.
Published: (2026)
OmniBooth: Learning Latent Control for Image Synthesis with Multi-modal Instruction
by: Li, Leheng, et al.
Published: (2024)
by: Li, Leheng, et al.
Published: (2024)
VaViM and VaVAM: Autonomous Driving through Video Generative Modeling
by: Bartoccioni, Florent, et al.
Published: (2025)
by: Bartoccioni, Florent, et al.
Published: (2025)
DyGeoVLN: Infusing Dynamic Geometry Foundation Model into Vision-Language Navigation
by: Liu, Xiangchen, et al.
Published: (2026)
by: Liu, Xiangchen, et al.
Published: (2026)
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
by: Lin, Minghui, et al.
Published: (2025)
by: Lin, Minghui, et al.
Published: (2025)
VAM-11: Master Equation for Particle Masses
by: Iskandarani, Omar
Published: (2025)
by: Iskandarani, Omar
Published: (2025)
Non-Negotiated Implicit ETSI VAM Clustering
by: Valle, Felipe, et al.
Published: (2025)
by: Valle, Felipe, et al.
Published: (2025)
CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
by: Song, Wenxuan, et al.
Published: (2026)
by: Song, Wenxuan, et al.
Published: (2026)
VAM: Verbalized Action Masking for Controllable Exploration in RL Post-Training -- A Chess Case Study
by: Zhang, Zhicheng, et al.
Published: (2026)
by: Zhang, Zhicheng, et al.
Published: (2026)
Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language Models
by: Xu, Tianshuo, et al.
Published: (2025)
by: Xu, Tianshuo, et al.
Published: (2025)
DisEnvisioner: Disentangled and Enriched Visual Prompt for Customized Image Generation
by: He, Jing, et al.
Published: (2024)
by: He, Jing, et al.
Published: (2024)
One‐Step Creation of a Spin‐Resonator‐Spin W State via Shortcut to Adiabaticity
by: Zhi‐Bo Feng, et al.
Published: (2026)
by: Zhi‐Bo Feng, et al.
Published: (2026)
RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data
by: Chen, Harold Haodong, et al.
Published: (2026)
by: Chen, Harold Haodong, et al.
Published: (2026)
Confidence-aware Self-Semantic Distillation on Knowledge Graph Embedding
by: Liu, Yichen, et al.
Published: (2022)
by: Liu, Yichen, et al.
Published: (2022)
PianoVAM: A Multimodal Piano Performance Dataset
by: Kim, Yonghyun, et al.
Published: (2025)
by: Kim, Yonghyun, et al.
Published: (2025)
Robust MIMO Semantic Communication with Imperfect CSI via Knowledge Distillation
by: Gong, Mingze, et al.
Published: (2025)
by: Gong, Mingze, et al.
Published: (2025)
Foresighted Online Policy Optimization with Interference
by: Xiang, Liner, et al.
Published: (2025)
by: Xiang, Liner, et al.
Published: (2025)
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
by: Fang, Luyang, et al.
Published: (2026)
by: Fang, Luyang, et al.
Published: (2026)
Semantic-Cohesive Knowledge Distillation for Deep Cross-modal Hashing
by: Sun, Changchang, et al.
Published: (2025)
by: Sun, Changchang, et al.
Published: (2025)
Distill Video Datasets into Images
by: Zhao, Zhenghao, et al.
Published: (2025)
by: Zhao, Zhenghao, et al.
Published: (2025)
Semantic-Guided Unsupervised Video Summarization
by: Liu, Haizhou, et al.
Published: (2026)
by: Liu, Haizhou, et al.
Published: (2026)
Delving Deep into Semantic Relation Distillation
by: Yan, Zhaoyi, et al.
Published: (2025)
by: Yan, Zhaoyi, et al.
Published: (2025)
RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark
by: Lei, Huashuo, et al.
Published: (2026)
by: Lei, Huashuo, et al.
Published: (2026)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
by: Chen, Jiayi, et al.
Published: (2025)
by: Chen, Jiayi, et al.
Published: (2025)
Multilateral Cascading Network for Semantic Segmentation of Large-Scale Outdoor Point Clouds
by: Gong, Haoran, et al.
Published: (2024)
by: Gong, Haoran, et al.
Published: (2024)
Similar Items
-
VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation
by: Zhong, Zhide, et al.
Published: (2026) -
DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models
by: Zhong, Zhide, et al.
Published: (2026) -
Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation
by: Yan, Haodong, et al.
Published: (2025) -
FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
by: Zhong, Zhide, et al.
Published: (2025) -
SyntheOcc: Synthesize Geometric-Controlled Street View Images through 3D Semantic MPIs
by: Li, Leheng, et al.
Published: (2024)