QuoVLA: Quotient Space for Vision-Language-Action Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xuan, Wu, Yinan, Duan, Haoran, Han, Jungong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
von: Li, Yixuan, et al.
Veröffentlicht: (2025)
von: Li, Yixuan, et al.
Veröffentlicht: (2025)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
THU-Warwick Submission for EPIC-KITCHEN Challenge 2025: Semi-Supervised Video Object Segmentation
von: Gao, Mingqi, et al.
Veröffentlicht: (2025)
von: Gao, Mingqi, et al.
Veröffentlicht: (2025)
Rethinking Score Distilling Sampling for 3D Editing and Generation
von: Miao, Xingyu, et al.
Veröffentlicht: (2025)
von: Miao, Xingyu, et al.
Veröffentlicht: (2025)
EvoVLA: Self-Evolving Vision-Language-Action Model
von: Liu, Zeting, et al.
Veröffentlicht: (2025)
von: Liu, Zeting, et al.
Veröffentlicht: (2025)
LLaDA-VLA: Vision Language Diffusion Action Models
von: Wen, Yuqing, et al.
Veröffentlicht: (2025)
von: Wen, Yuqing, et al.
Veröffentlicht: (2025)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval
von: Shen, Leqi, et al.
Veröffentlicht: (2025)
von: Shen, Leqi, et al.
Veröffentlicht: (2025)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
von: Ye, Angen, et al.
Veröffentlicht: (2025)
von: Ye, Angen, et al.
Veröffentlicht: (2025)
Vision-Language-Action (VLA) Models: Concepts, Progress, Applications and Challenges
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models
von: Yang, Yantai, et al.
Veröffentlicht: (2025)
von: Yang, Yantai, et al.
Veröffentlicht: (2025)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
von: Li, Runhao, et al.
Veröffentlicht: (2025)
von: Li, Runhao, et al.
Veröffentlicht: (2025)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving
von: Luo, Yuechen, et al.
Veröffentlicht: (2026)
von: Luo, Yuechen, et al.
Veröffentlicht: (2026)
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
von: Ye, Jiacheng, et al.
Veröffentlicht: (2025)
von: Ye, Jiacheng, et al.
Veröffentlicht: (2025)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
von: jia, Feiyang, et al.
Veröffentlicht: (2026)
von: jia, Feiyang, et al.
Veröffentlicht: (2026)
Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models
von: Chi, Haohan, et al.
Veröffentlicht: (2025)
von: Chi, Haohan, et al.
Veröffentlicht: (2025)
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2025)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2025)
MAIN-VLA: Modeling Abstraction of Intention and eNvironment for Vision-Language-Action Models
von: Zhou, Zheyuan, et al.
Veröffentlicht: (2026)
von: Zhou, Zheyuan, et al.
Veröffentlicht: (2026)
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
von: Zheng, Jinliang, et al.
Veröffentlicht: (2025)
von: Zheng, Jinliang, et al.
Veröffentlicht: (2025)
VLA-IAP: Training-Free Visual Token Pruning via Interaction Alignment for Vision-Language-Action Models
von: Cheng, Jintao, et al.
Veröffentlicht: (2026)
von: Cheng, Jintao, et al.
Veröffentlicht: (2026)
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
von: Sun, Jingwen, et al.
Veröffentlicht: (2026)
von: Sun, Jingwen, et al.
Veröffentlicht: (2026)
VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
von: Dong, Shaoqi, et al.
Veröffentlicht: (2025)
von: Dong, Shaoqi, et al.
Veröffentlicht: (2025)
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
von: Wang, Yating, et al.
Veröffentlicht: (2025)
von: Wang, Yating, et al.
Veröffentlicht: (2025)
UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models
von: Zhang, Qiyao, et al.
Veröffentlicht: (2026)
von: Zhang, Qiyao, et al.
Veröffentlicht: (2026)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
von: Chen, Jiayi, et al.
Veröffentlicht: (2025)
von: Chen, Jiayi, et al.
Veröffentlicht: (2025)
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
von: Du, Fan, et al.
Veröffentlicht: (2026)
von: Du, Fan, et al.
Veröffentlicht: (2026)
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
von: Yuan, Tianyuan, et al.
Veröffentlicht: (2025)
von: Yuan, Tianyuan, et al.
Veröffentlicht: (2025)
GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning
von: Ma, Guoqing, et al.
Veröffentlicht: (2026)
von: Ma, Guoqing, et al.
Veröffentlicht: (2026)
CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving
von: Arai, Hidehisa, et al.
Veröffentlicht: (2024)
von: Arai, Hidehisa, et al.
Veröffentlicht: (2024)
StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early Observation
von: Shi, Yiran, et al.
Veröffentlicht: (2026)
von: Shi, Yiran, et al.
Veröffentlicht: (2026)
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
von: Zhang, Wenyao, et al.
Veröffentlicht: (2025)
von: Zhang, Wenyao, et al.
Veröffentlicht: (2025)
SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model
von: Zhou, Zewei, et al.
Veröffentlicht: (2026)
von: Zhou, Zewei, et al.
Veröffentlicht: (2026)
SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving
von: You, Zihan, et al.
Veröffentlicht: (2026)
von: You, Zihan, et al.
Veröffentlicht: (2026)
FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation
von: Zhao, Ruiteng, et al.
Veröffentlicht: (2026)
von: Zhao, Ruiteng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
von: Wang, Xin, et al.
Veröffentlicht: (2025) -
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023) -
QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
von: Li, Yixuan, et al.
Veröffentlicht: (2025) -
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025) -
THU-Warwick Submission for EPIC-KITCHEN Challenge 2025: Semi-Supervised Video Object Segmentation
von: Gao, Mingqi, et al.
Veröffentlicht: (2025)