Words to Wheels: Vision-Based Autonomous Driving Understanding Human Language Instructions Using Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ryu, Chanhoe, Seong, Hyunki, Lee, Daegyu, Moon, Seongwoo, Min, Sungjae, Shim, D. Hyunchul |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing State Estimator for Autonomous Racing : Leveraging Multi-modal System and Managing Computing Resources
by: Lee, Daegyu, et al.
Published: (2023)
by: Lee, Daegyu, et al.
Published: (2023)
A Versatile Door Opening System with Mobile Manipulator through Adaptive Position-Force Control and Reinforcement Learning
by: Kang, Gyuree, et al.
Published: (2023)
by: Kang, Gyuree, et al.
Published: (2023)
VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
by: Seong, Hyunki, et al.
Published: (2025)
by: Seong, Hyunki, et al.
Published: (2025)
TempFuser: Learning Agile, Tactical, and Acrobatic Flight Maneuvers Using a Long Short-Term Temporal Fusion Transformer
by: Seong, Hyunki, et al.
Published: (2023)
by: Seong, Hyunki, et al.
Published: (2023)
Skill Q-Network: Learning Adaptive Skill Ensemble for Mapless Navigation in Unknown Environments
by: Seong, Hyunki, et al.
Published: (2024)
by: Seong, Hyunki, et al.
Published: (2024)
Learning from Demonstration with Hierarchical Policy Abstractions Toward High-Performance and Courteous Autonomous Racing
by: Chung, Chanyoung, et al.
Published: (2024)
by: Chung, Chanyoung, et al.
Published: (2024)
Self-Supervised Interpretable End-to-End Learning via Latent Functional Modularity
by: Seong, Hyunki, et al.
Published: (2024)
by: Seong, Hyunki, et al.
Published: (2024)
Antagonistic Bowden-Cable Actuation of a Lightweight Robotic Hand: Toward Dexterous Manipulation for Payload Constrained Humanoids
by: Min, Sungjae, et al.
Published: (2025)
by: Min, Sungjae, et al.
Published: (2025)
MonoDINO-DETR: Depth-Enhanced Monocular 3D Object Detection Using a Vision Foundation Model
by: Kim, Jihyeok, et al.
Published: (2025)
by: Kim, Jihyeok, et al.
Published: (2025)
From Words to Wheels: Automated Style-Customized Policy Generation for Autonomous Driving
by: Han, Xu, et al.
Published: (2024)
by: Han, Xu, et al.
Published: (2024)
LLM-Flax : Generalizable Robotic Task Planning via Neuro-Symbolic Approaches with Large Language Models
by: Kim, Seongmin, et al.
Published: (2026)
by: Kim, Seongmin, et al.
Published: (2026)
Robust Tightly-Coupled Filter-Based Monocular Visual-Inertial State Estimation and Graph-Based Evaluation for Autonomous Drone Racing
by: Azhari, Maulana Bisyir, et al.
Published: (2026)
by: Azhari, Maulana Bisyir, et al.
Published: (2026)
DINO-VO: A Feature-based Visual Odometry Leveraging a Visual Foundation Model
by: Azhari, Maulana Bisyir, et al.
Published: (2025)
by: Azhari, Maulana Bisyir, et al.
Published: (2025)
SUPER-AD: Semantic Uncertainty-aware Planning for End-to-End Robust Autonomous Driving
by: Ryu, Wonjeong, et al.
Published: (2025)
by: Ryu, Wonjeong, et al.
Published: (2025)
GraspCorrect: Robotic Grasp Correction via Vision-Language Model-Guided Feedback
by: Lee, Sungjae, et al.
Published: (2025)
by: Lee, Sungjae, et al.
Published: (2025)
A Collaborative Team of UAV-Hexapod for an Autonomous Retrieval System in GNSS-Denied Maritime Environments
by: Lee, Seungwook, et al.
Published: (2024)
by: Lee, Seungwook, et al.
Published: (2024)
Drift-Corrected Monocular VIO and Perception-Aware Planning for Autonomous Drone Racing
by: Azhari, Maulana Bisyir, et al.
Published: (2025)
by: Azhari, Maulana Bisyir, et al.
Published: (2025)
Miniature Testbed for Validating Multi-Agent Cooperative Autonomous Driving
by: Bae, Hyunchul, et al.
Published: (2025)
by: Bae, Hyunchul, et al.
Published: (2025)
Post-Training and Test-Time Scaling of Generative Agent Behavior Models for Interactive Autonomous Driving
by: Seong, Hyunki, et al.
Published: (2025)
by: Seong, Hyunki, et al.
Published: (2025)
Wheel Odometry-Based Localization for Autonomous Wheelchair
by: Paryanto, P, et al.
Published: (2023)
by: Paryanto, P, et al.
Published: (2023)
Don't Shake the Wheel: Momentum-Aware Planning in End-to-End Autonomous Driving
by: Song, Ziying, et al.
Published: (2025)
by: Song, Ziying, et al.
Published: (2025)
MORPH Wheel: A Passive Variable-Radius Wheel Embedding Mechanical Behavior Logic for Input-Responsive Transformation
by: Jang, JaeHyung, et al.
Published: (2026)
by: Jang, JaeHyung, et al.
Published: (2026)
VCA: Vision-Click-Action Framework for Precise Manipulation of Segmented Objects in Target Ambiguous Environments
by: Kim, Donggeon, et al.
Published: (2026)
by: Kim, Donggeon, et al.
Published: (2026)
VLM-MPC: Vision Language Foundation Model (VLM)-Guided Model Predictive Controller (MPC) for Autonomous Driving
by: Long, Keke, et al.
Published: (2024)
by: Long, Keke, et al.
Published: (2024)
NDST: Neural Driving Style Transfer for Human-Like Vision-Based Autonomous Driving
by: Kim, Donghyun, et al.
Published: (2024)
by: Kim, Donghyun, et al.
Published: (2024)
Learning to Drift with Individual Wheel Drive: Maneuvering Autonomous Vehicle at the Handling Limits
by: Zhou, Yihan, et al.
Published: (2025)
by: Zhou, Yihan, et al.
Published: (2025)
SPIBOT: A Drone-Tethered Mobile Gripper for Robust Aerial Object Retrieval in Dynamic Environments
by: Kang, Gyuree, et al.
Published: (2024)
by: Kang, Gyuree, et al.
Published: (2024)
A Vision-Language-Action Model with Visual Prompt for OFF-Road Autonomous Driving
by: Zhang, Liangdong, et al.
Published: (2026)
by: Zhang, Liangdong, et al.
Published: (2026)
World Modeling for Autonomous Wheel Loaders
by: Aoshima, Koji, et al.
Published: (2023)
by: Aoshima, Koji, et al.
Published: (2023)
ELiOT: End‐to‐end LiDAR odometry with transformers harnessing real‐world, simulated, and digital twin
by: Daegyu Lee, et al.
Published: (2025)
by: Daegyu Lee, et al.
Published: (2025)
StyleVLA: Driving Style-Aware Vision Language Action Model for Autonomous Driving
by: Gao, Yuan, et al.
Published: (2026)
by: Gao, Yuan, et al.
Published: (2026)
Integration of Computer Vision with Adaptive Control for Autonomous Driving Using ADORE
by: Ahammed, Abu Shad, et al.
Published: (2025)
by: Ahammed, Abu Shad, et al.
Published: (2025)
Multimodal Human-Autonomous Agents Interaction Using Pre-Trained Language and Visual Foundation Models
by: Nwankwo, Linus, et al.
Published: (2024)
by: Nwankwo, Linus, et al.
Published: (2024)
Natural Language Instructions for Scene-Responsive Human-in-the-Loop Motion Planning in Autonomous Driving using Vision-Language-Action Models
by: Martinez-Sanchez, Angel, et al.
Published: (2026)
by: Martinez-Sanchez, Angel, et al.
Published: (2026)
Bench2Drive-VL: Benchmarks for Closed-Loop Autonomous Driving with Vision-Language Models
by: Jia, Xiaosong, et al.
Published: (2026)
by: Jia, Xiaosong, et al.
Published: (2026)
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
by: Schäfer, Finn Rasmus, et al.
Published: (2026)
by: Schäfer, Finn Rasmus, et al.
Published: (2026)
Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
by: Hu, Tianshuai, et al.
Published: (2025)
by: Hu, Tianshuai, et al.
Published: (2025)
Autonomous Improvement of Instruction Following Skills via Foundation Models
by: Zhou, Zhiyuan, et al.
Published: (2024)
by: Zhou, Zhiyuan, et al.
Published: (2024)
Words2Contact: Identifying Support Contacts from Verbal Instructions Using Foundation Models
by: Totsila, Dionis, et al.
Published: (2024)
by: Totsila, Dionis, et al.
Published: (2024)
Causality-Aware End-to-End Autonomous Driving via Ego-Centric Joint Scene Modeling
by: Moon, Seokha, et al.
Published: (2026)
by: Moon, Seokha, et al.
Published: (2026)
Similar Items
-
Enhancing State Estimator for Autonomous Racing : Leveraging Multi-modal System and Managing Computing Resources
by: Lee, Daegyu, et al.
Published: (2023) -
A Versatile Door Opening System with Mobile Manipulator through Adaptive Position-Force Control and Reinforcement Learning
by: Kang, Gyuree, et al.
Published: (2023) -
VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
by: Seong, Hyunki, et al.
Published: (2025) -
TempFuser: Learning Agile, Tactical, and Acrobatic Flight Maneuvers Using a Long Short-Term Temporal Fusion Transformer
by: Seong, Hyunki, et al.
Published: (2023) -
Skill Q-Network: Learning Adaptive Skill Ensemble for Mapless Navigation in Unknown Environments
by: Seong, Hyunki, et al.
Published: (2024)