Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation
Fuente:
arXiv
Guardado en:
| Autores principales: | Shi, Jin, Zhang, Brady, Lu, Yishun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
por: Du, Fan, et al.
Publicado: (2026)
por: Du, Fan, et al.
Publicado: (2026)
Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models
por: Lu, Yishun, et al.
Publicado: (2026)
por: Lu, Yishun, et al.
Publicado: (2026)
TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
por: Zhang, Shu-Hao, et al.
Publicado: (2025)
por: Zhang, Shu-Hao, et al.
Publicado: (2025)
Multimodal Distribution Matching for Vision-Language Dataset Distillation
por: Jeong, Jongoh, et al.
Publicado: (2026)
por: Jeong, Jongoh, et al.
Publicado: (2026)
VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation
por: Xu, Changhua, et al.
Publicado: (2026)
por: Xu, Changhua, et al.
Publicado: (2026)
Remodeling Semantic Relationships in Vision-Language Fine-Tuning
por: Wu, Xiangyang, et al.
Publicado: (2025)
por: Wu, Xiangyang, et al.
Publicado: (2025)
When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
por: Yan, Yuping, et al.
Publicado: (2025)
por: Yan, Yuping, et al.
Publicado: (2025)
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
por: Gao, Mingjian, et al.
Publicado: (2026)
por: Gao, Mingjian, et al.
Publicado: (2026)
INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling
por: Dong, Xin, et al.
Publicado: (2025)
por: Dong, Xin, et al.
Publicado: (2025)
VITA: Vision-to-Action Flow Matching Policy
por: Gao, Dechen, et al.
Publicado: (2025)
por: Gao, Dechen, et al.
Publicado: (2025)
ARFlow: Human Action-Reaction Flow Matching with Physical Guidance
por: Jiang, Wentao, et al.
Publicado: (2025)
por: Jiang, Wentao, et al.
Publicado: (2025)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
por: Zhang, Jianke, et al.
Publicado: (2026)
por: Zhang, Jianke, et al.
Publicado: (2026)
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
por: Chen, Xinyi, et al.
Publicado: (2025)
por: Chen, Xinyi, et al.
Publicado: (2025)
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
por: Xu, Chenkai, et al.
Publicado: (2025)
por: Xu, Chenkai, et al.
Publicado: (2025)
Adversarial Prompt Distillation for Vision-Language Models
por: Luo, Lin, et al.
Publicado: (2024)
por: Luo, Lin, et al.
Publicado: (2024)
CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games
por: Chen, Peng, et al.
Publicado: (2025)
por: Chen, Peng, et al.
Publicado: (2025)
MAGIC: Meta-Ability Guided Interactive Chain-of-Distillation for Effective-and-Efficient Vision-and-Language Navigation
por: Wang, Liuyi, et al.
Publicado: (2024)
por: Wang, Liuyi, et al.
Publicado: (2024)
DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models
por: Wang, Qichao, et al.
Publicado: (2026)
por: Wang, Qichao, et al.
Publicado: (2026)
Representation Calibration and Uncertainty Guidance for Class-Incremental Learning based on Vision Language Model
por: Tan, Jiantao, et al.
Publicado: (2025)
por: Tan, Jiantao, et al.
Publicado: (2025)
Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
por: Qi, Xiuxiu, et al.
Publicado: (2025)
por: Qi, Xiuxiu, et al.
Publicado: (2025)
Information-Theoretic Constraints for Continual Vision-Language-Action Alignment
por: Zhao, Libang, et al.
Publicado: (2026)
por: Zhao, Libang, et al.
Publicado: (2026)
X-Distill: Cross-Architecture Vision Distillation for Visuomotor Learning
por: Shao, Maanping, et al.
Publicado: (2026)
por: Shao, Maanping, et al.
Publicado: (2026)
IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model
por: Jiang, Anqing, et al.
Publicado: (2025)
por: Jiang, Anqing, et al.
Publicado: (2025)
NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning
por: Rawal, Ishaan, et al.
Publicado: (2026)
por: Rawal, Ishaan, et al.
Publicado: (2026)
Dual-Model Distillation for Efficient Action Classification with Hybrid Edge-Cloud Solution
por: Wei, Timothy, et al.
Publicado: (2024)
por: Wei, Timothy, et al.
Publicado: (2024)
When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models
por: Lu, Hui, et al.
Publicado: (2025)
por: Lu, Hui, et al.
Publicado: (2025)
A Survey on Efficient Vision-Language-Action Models
por: Yu, Zhaoshu, et al.
Publicado: (2025)
por: Yu, Zhaoshu, et al.
Publicado: (2025)
Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
por: Yang, Yi, et al.
Publicado: (2025)
por: Yang, Yi, et al.
Publicado: (2025)
Single-Sample Black-Box Membership Inference Attack against Vision-Language Models via Cross-modal Semantic Alignment
por: Li, Jiaqing, et al.
Publicado: (2026)
por: Li, Jiaqing, et al.
Publicado: (2026)
SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
por: Ma, Qiwei, et al.
Publicado: (2025)
por: Ma, Qiwei, et al.
Publicado: (2025)
Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation
por: Yang, Yang, et al.
Publicado: (2024)
por: Yang, Yang, et al.
Publicado: (2024)
InkSight: Offline-to-Online Handwriting Conversion by Teaching Vision-Language Models to Read and Write
por: Mitrevski, Blagoj, et al.
Publicado: (2024)
por: Mitrevski, Blagoj, et al.
Publicado: (2024)
Active Multimodal Distillation for Few-shot Action Recognition
por: Feng, Weijia, et al.
Publicado: (2025)
por: Feng, Weijia, et al.
Publicado: (2025)
Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?
por: He, Jingtao, et al.
Publicado: (2026)
por: He, Jingtao, et al.
Publicado: (2026)
D3S2: Diffusion-Guided Dataset Distillation for Semantic Segmentation
por: Zheng, Wenjie, et al.
Publicado: (2026)
por: Zheng, Wenjie, et al.
Publicado: (2026)
Survey on Vision-Language-Action Models
por: Adilkhanov, Adilzhan, et al.
Publicado: (2025)
por: Adilkhanov, Adilzhan, et al.
Publicado: (2025)
Enhancing Medical Large Vision-Language Models via Alignment Distillation
por: Chang, Aofei, et al.
Publicado: (2025)
por: Chang, Aofei, et al.
Publicado: (2025)
AUFormer: Vision Transformers are Parameter-Efficient Facial Action Unit Detectors
por: Yuan, Kaishen, et al.
Publicado: (2024)
por: Yuan, Kaishen, et al.
Publicado: (2024)
Attention-Guided Patch-Wise Sparse Adversarial Attacks on Vision-Language-Action Models
por: Zhang, Naifu, et al.
Publicado: (2025)
por: Zhang, Naifu, et al.
Publicado: (2025)
Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual Guidance
por: Chen, Xinrong, et al.
Publicado: (2026)
por: Chen, Xinrong, et al.
Publicado: (2026)
Ejemplares similares
-
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
por: Du, Fan, et al.
Publicado: (2026) -
Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models
por: Lu, Yishun, et al.
Publicado: (2026) -
TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
por: Zhang, Shu-Hao, et al.
Publicado: (2025) -
Multimodal Distribution Matching for Vision-Language Dataset Distillation
por: Jeong, Jongoh, et al.
Publicado: (2026) -
VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation
por: Xu, Changhua, et al.
Publicado: (2026)