Saved in:
| Main Authors: | Shi, Jin, Zhang, Brady, Lu, Yishun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.16241 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models
by: Lu, Yishun, et al.
Published: (2026)
by: Lu, Yishun, et al.
Published: (2026)
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
by: Du, Fan, et al.
Published: (2026)
by: Du, Fan, et al.
Published: (2026)
Multimodal Distribution Matching for Vision-Language Dataset Distillation
by: Jeong, Jongoh, et al.
Published: (2026)
by: Jeong, Jongoh, et al.
Published: (2026)
TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
by: Zhang, Shu-Hao, et al.
Published: (2025)
by: Zhang, Shu-Hao, et al.
Published: (2025)
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
by: Xu, Chenkai, et al.
Published: (2025)
by: Xu, Chenkai, et al.
Published: (2025)
VITA: Vision-to-Action Flow Matching Policy
by: Gao, Dechen, et al.
Published: (2025)
by: Gao, Dechen, et al.
Published: (2025)
VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation
by: Xu, Changhua, et al.
Published: (2026)
by: Xu, Changhua, et al.
Published: (2026)
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
by: Chen, Xinyi, et al.
Published: (2025)
by: Chen, Xinyi, et al.
Published: (2025)
When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
by: Yan, Yuping, et al.
Published: (2025)
by: Yan, Yuping, et al.
Published: (2025)
ARFlow: Human Action-Reaction Flow Matching with Physical Guidance
by: Jiang, Wentao, et al.
Published: (2025)
by: Jiang, Wentao, et al.
Published: (2025)
Remodeling Semantic Relationships in Vision-Language Fine-Tuning
by: Wu, Xiangyang, et al.
Published: (2025)
by: Wu, Xiangyang, et al.
Published: (2025)
INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling
by: Dong, Xin, et al.
Published: (2025)
by: Dong, Xin, et al.
Published: (2025)
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
by: Gao, Mingjian, et al.
Published: (2026)
by: Gao, Mingjian, et al.
Published: (2026)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
by: Zhang, Jianke, et al.
Published: (2026)
by: Zhang, Jianke, et al.
Published: (2026)
Adversarial Prompt Distillation for Vision-Language Models
by: Luo, Lin, et al.
Published: (2024)
by: Luo, Lin, et al.
Published: (2024)
A Survey on Efficient Vision-Language-Action Models
by: Yu, Zhaoshu, et al.
Published: (2025)
by: Yu, Zhaoshu, et al.
Published: (2025)
Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
by: Qi, Xiuxiu, et al.
Published: (2025)
by: Qi, Xiuxiu, et al.
Published: (2025)
CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games
by: Chen, Peng, et al.
Published: (2025)
by: Chen, Peng, et al.
Published: (2025)
DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models
by: Wang, Qichao, et al.
Published: (2026)
by: Wang, Qichao, et al.
Published: (2026)
IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model
by: Jiang, Anqing, et al.
Published: (2025)
by: Jiang, Anqing, et al.
Published: (2025)
MAGIC: Meta-Ability Guided Interactive Chain-of-Distillation for Effective-and-Efficient Vision-and-Language Navigation
by: Wang, Liuyi, et al.
Published: (2024)
by: Wang, Liuyi, et al.
Published: (2024)
Representation Calibration and Uncertainty Guidance for Class-Incremental Learning based on Vision Language Model
by: Tan, Jiantao, et al.
Published: (2025)
by: Tan, Jiantao, et al.
Published: (2025)
Information-Theoretic Constraints for Continual Vision-Language-Action Alignment
by: Zhao, Libang, et al.
Published: (2026)
by: Zhao, Libang, et al.
Published: (2026)
X-Distill: Cross-Architecture Vision Distillation for Visuomotor Learning
by: Shao, Maanping, et al.
Published: (2026)
by: Shao, Maanping, et al.
Published: (2026)
NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning
by: Rawal, Ishaan, et al.
Published: (2026)
by: Rawal, Ishaan, et al.
Published: (2026)
Dual-Model Distillation for Efficient Action Classification with Hybrid Edge-Cloud Solution
by: Wei, Timothy, et al.
Published: (2024)
by: Wei, Timothy, et al.
Published: (2024)
Spatial Transcriptomics as Images for Large-Scale Pretraining
by: Zhu, Yishun, et al.
Published: (2026)
by: Zhu, Yishun, et al.
Published: (2026)
When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models
by: Lu, Hui, et al.
Published: (2025)
by: Lu, Hui, et al.
Published: (2025)
Single-Sample Black-Box Membership Inference Attack against Vision-Language Models via Cross-modal Semantic Alignment
by: Li, Jiaqing, et al.
Published: (2026)
by: Li, Jiaqing, et al.
Published: (2026)
Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
by: Ma, Qiwei, et al.
Published: (2025)
by: Ma, Qiwei, et al.
Published: (2025)
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
by: Li, Bangzheng, et al.
Published: (2025)
by: Li, Bangzheng, et al.
Published: (2025)
Survey on Vision-Language-Action Models
by: Adilkhanov, Adilzhan, et al.
Published: (2025)
by: Adilkhanov, Adilzhan, et al.
Published: (2025)
Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation
by: Yang, Yang, et al.
Published: (2024)
by: Yang, Yang, et al.
Published: (2024)
InkSight: Offline-to-Online Handwriting Conversion by Teaching Vision-Language Models to Read and Write
by: Mitrevski, Blagoj, et al.
Published: (2024)
by: Mitrevski, Blagoj, et al.
Published: (2024)
Active Multimodal Distillation for Few-shot Action Recognition
by: Feng, Weijia, et al.
Published: (2025)
by: Feng, Weijia, et al.
Published: (2025)
PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation
by: Fan, Qingyu, et al.
Published: (2026)
by: Fan, Qingyu, et al.
Published: (2026)
Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?
by: He, Jingtao, et al.
Published: (2026)
by: He, Jingtao, et al.
Published: (2026)
LIBERO-X: Robustness Litmus for Vision-Language-Action Models
by: Wang, Guodong, et al.
Published: (2026)
by: Wang, Guodong, et al.
Published: (2026)
D3S2: Diffusion-Guided Dataset Distillation for Semantic Segmentation
by: Zheng, Wenjie, et al.
Published: (2026)
by: Zheng, Wenjie, et al.
Published: (2026)
Similar Items
-
Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models
by: Lu, Yishun, et al.
Published: (2026) -
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
by: Du, Fan, et al.
Published: (2026) -
Multimodal Distribution Matching for Vision-Language Dataset Distillation
by: Jeong, Jongoh, et al.
Published: (2026) -
TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
by: Zhang, Shu-Hao, et al.
Published: (2025) -
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
by: Xu, Chenkai, et al.
Published: (2025)