NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rawal, Ishaan, Gupta, Shubh, Hu, Yihan, Zhan, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interactive Video Generation via Domain Adaptation
von: Rawal, Ishaan, et al.
Veröffentlicht: (2025)
von: Rawal, Ishaan, et al.
Veröffentlicht: (2025)
VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
von: Rawal, Ishaan Singh, et al.
Veröffentlicht: (2023)
von: Rawal, Ishaan Singh, et al.
Veröffentlicht: (2023)
Learning Vision-Language-Action World Models for Autonomous Driving
von: Wang, Guoqing, et al.
Veröffentlicht: (2026)
von: Wang, Guoqing, et al.
Veröffentlicht: (2026)
A Survey on Vision-Language-Action Models for Autonomous Driving
von: Jiang, Sicong, et al.
Veröffentlicht: (2025)
von: Jiang, Sicong, et al.
Veröffentlicht: (2025)
CARE Drive A Framework for Evaluating Reason-Responsiveness of Vision Language Models in Automated Driving
von: Suryana, Lucas Elbert, et al.
Veröffentlicht: (2026)
von: Suryana, Lucas Elbert, et al.
Veröffentlicht: (2026)
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
von: Mei, Xiaodong, et al.
Veröffentlicht: (2026)
von: Mei, Xiaodong, et al.
Veröffentlicht: (2026)
ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations
von: Zhou, Yuhao, et al.
Veröffentlicht: (2026)
von: Zhou, Yuhao, et al.
Veröffentlicht: (2026)
Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
von: Yang, Yi, et al.
Veröffentlicht: (2025)
von: Yang, Yi, et al.
Veröffentlicht: (2025)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
ATA: Bridging Implicit Reasoning with Attention-Guided and Action-Guided Inference for Vision-Language Action Models
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis
von: Zhang, Congzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Congzhi, et al.
Veröffentlicht: (2025)
RAD: Retrieval-Augmented Decision-Making of Meta-Actions with Vision-Language Models in Autonomous Driving
von: Wang, Yujin, et al.
Veröffentlicht: (2025)
von: Wang, Yujin, et al.
Veröffentlicht: (2025)
DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
von: Yang, Zhenjie, et al.
Veröffentlicht: (2025)
von: Yang, Zhenjie, et al.
Veröffentlicht: (2025)
E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving
von: Tang, Yihong, et al.
Veröffentlicht: (2025)
von: Tang, Yihong, et al.
Veröffentlicht: (2025)
HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios
von: Wang, Daming, et al.
Veröffentlicht: (2025)
von: Wang, Daming, et al.
Veröffentlicht: (2025)
MedSAGa: Few-shot Memory Efficient Medical Image Segmentation using Gradient Low-Rank Projection in SAM
von: Mahla, Navyansh, et al.
Veröffentlicht: (2024)
von: Mahla, Navyansh, et al.
Veröffentlicht: (2024)
Prompting Large Vision-Language Models for Compositional Reasoning
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024)
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024)
Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?
von: He, Jingtao, et al.
Veröffentlicht: (2026)
von: He, Jingtao, et al.
Veröffentlicht: (2026)
Vision Language Models in Autonomous Driving: A Survey and Outlook
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2023)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2023)
Enhancing Vision-Language Models for Autonomous Driving through Task-Specific Prompting and Spatial Reasoning
von: Wu, Aodi, et al.
Veröffentlicht: (2025)
von: Wu, Aodi, et al.
Veröffentlicht: (2025)
Less is More: Lean yet Powerful Vision-Language Model for Autonomous Driving
von: Yang, Sheng, et al.
Veröffentlicht: (2025)
von: Yang, Sheng, et al.
Veröffentlicht: (2025)
DriveAction: A Benchmark for Exploring Human-like Driving Decisions in VLA Models
von: Hao, Yuhan, et al.
Veröffentlicht: (2025)
von: Hao, Yuhan, et al.
Veröffentlicht: (2025)
Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering in Autonomous Driving
von: Gopalkrishnan, Akshay, et al.
Veröffentlicht: (2024)
von: Gopalkrishnan, Akshay, et al.
Veröffentlicht: (2024)
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
von: Du, Fan, et al.
Veröffentlicht: (2026)
von: Du, Fan, et al.
Veröffentlicht: (2026)
SimpleLLM4AD: An End-to-End Vision-Language Model with Graph Visual Question Answering for Autonomous Driving
von: Zheng, Peiru, et al.
Veröffentlicht: (2024)
von: Zheng, Peiru, et al.
Veröffentlicht: (2024)
CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games
von: Chen, Peng, et al.
Veröffentlicht: (2025)
von: Chen, Peng, et al.
Veröffentlicht: (2025)
Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation
von: Shi, Jin, et al.
Veröffentlicht: (2026)
von: Shi, Jin, et al.
Veröffentlicht: (2026)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
von: Tian, Kexin, et al.
Veröffentlicht: (2025)
von: Tian, Kexin, et al.
Veröffentlicht: (2025)
VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
von: Ogezi, Michael, et al.
Veröffentlicht: (2025)
von: Ogezi, Michael, et al.
Veröffentlicht: (2025)
DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale
von: Zuo, Sicheng, et al.
Veröffentlicht: (2026)
von: Zuo, Sicheng, et al.
Veröffentlicht: (2026)
Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving
von: Jiang, Hao, et al.
Veröffentlicht: (2025)
von: Jiang, Hao, et al.
Veröffentlicht: (2025)
AD-EE: Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving
von: Huang, Lianming, et al.
Veröffentlicht: (2025)
von: Huang, Lianming, et al.
Veröffentlicht: (2025)
Agentic Surgical AI: Surgeon Style Fingerprinting and Privacy Risk Quantification via Discrete Diffusion in a Vision-Language-Action Framework
von: Zhan, Huixin, et al.
Veröffentlicht: (2025)
von: Zhan, Huixin, et al.
Veröffentlicht: (2025)
Survey on Vision-Language-Action Models
von: Adilkhanov, Adilzhan, et al.
Veröffentlicht: (2025)
von: Adilkhanov, Adilzhan, et al.
Veröffentlicht: (2025)
DeeAD: Dynamic Early Exit of Vision-Language Action for Efficient Autonomous Driving
von: HU, Haibo, et al.
Veröffentlicht: (2025)
von: HU, Haibo, et al.
Veröffentlicht: (2025)
NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models
von: Pandya, Pranshu, et al.
Veröffentlicht: (2024)
von: Pandya, Pranshu, et al.
Veröffentlicht: (2024)
Conformal Predictions for Human Action Recognition with Vision-Language Models
von: Tim, Bary, et al.
Veröffentlicht: (2025)
von: Tim, Bary, et al.
Veröffentlicht: (2025)
Efficient Driving Behavior Narration and Reasoning on Edge Device Using Large Language Models
von: Huang, Yizhou, et al.
Veröffentlicht: (2024)
von: Huang, Yizhou, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Interactive Video Generation via Domain Adaptation
von: Rawal, Ishaan, et al.
Veröffentlicht: (2025) -
VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving
von: Zhao, Rui, et al.
Veröffentlicht: (2026) -
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
von: Rawal, Ishaan Singh, et al.
Veröffentlicht: (2023) -
Learning Vision-Language-Action World Models for Autonomous Driving
von: Wang, Guoqing, et al.
Veröffentlicht: (2026) -
A Survey on Vision-Language-Action Models for Autonomous Driving
von: Jiang, Sicong, et al.
Veröffentlicht: (2025)