VLM-Auto: VLM-based Autonomous Driving Assistant with Human-like Behavior and Understanding for Complex Road Scenes
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Ziang, Yagudin, Zakhar, Lykov, Artem, Konenkov, Mikhail, Tsetserukou, Dzmitry |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VDT-Auto: End-to-end Autonomous Driving with VLM-Guided Diffusion Transformers
by: Guo, Ziang, et al.
Published: (2025)
by: Guo, Ziang, et al.
Published: (2025)
FADet: A Multi-sensor 3D Object Detection Network based on Local Featured Attention
by: Guo, Ziang, et al.
Published: (2024)
by: Guo, Ziang, et al.
Published: (2024)
METDrive: Multi-modal End-to-end Autonomous Driving with Temporal Guidance
by: Guo, Ziang, et al.
Published: (2024)
by: Guo, Ziang, et al.
Published: (2024)
HawkDrive: A Transformer-driven Visual Perception System for Autonomous Driving in Night Scene
by: Guo, Ziang, et al.
Published: (2024)
by: Guo, Ziang, et al.
Published: (2024)
VR-GPT: Visual Language Model for Intelligent Virtual Reality Applications
by: Konenkov, Mikhail, et al.
Published: (2024)
by: Konenkov, Mikhail, et al.
Published: (2024)
SwarmVLM: VLM-Guided Impedance Control for Autonomous Navigation of Heterogeneous Robots in Dynamic Warehousing
by: Zafar, Malaika, et al.
Published: (2025)
by: Zafar, Malaika, et al.
Published: (2025)
GestLLM: Advanced Hand Gesture Interpretation via Large Language Models for Human-Robot Interaction
by: Kobzarev, Oleg, et al.
Published: (2025)
by: Kobzarev, Oleg, et al.
Published: (2025)
FlockGPT: Guiding UAV Flocking with Linguistic Orchestration
by: Lykov, Artem, et al.
Published: (2024)
by: Lykov, Artem, et al.
Published: (2024)
GestOS: Advanced Hand Gesture Interpretation via Large Language Models to control Any Type of Robot
by: Lykov, Artem, et al.
Published: (2025)
by: Lykov, Artem, et al.
Published: (2025)
CognitiveDog: Large Multimodal Model Based System to Translate Vision and Language into Action of Quadruped Robot
by: Lykov, Artem, et al.
Published: (2024)
by: Lykov, Artem, et al.
Published: (2024)
CognitiveOS: Large Multimodal Model based System to Endow Any Type of Robot with Generative AI
by: Lykov, Artem, et al.
Published: (2024)
by: Lykov, Artem, et al.
Published: (2024)
SafeHumanoid: VLM-RAG-driven Control of Upper Body Impedance for Humanoid Robot
by: Mahmoud, Yara, et al.
Published: (2025)
by: Mahmoud, Yara, et al.
Published: (2025)
Robots Can Feel: LLM-based Framework for Robot Ethical Reasoning
by: Lykov, Artem, et al.
Published: (2024)
by: Lykov, Artem, et al.
Published: (2024)
RaceVLA: VLA-based Racing Drone Navigation with Human-like Behaviour
by: Serpiva, Valerii, et al.
Published: (2025)
by: Serpiva, Valerii, et al.
Published: (2025)
PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models
by: Lykov, Artem, et al.
Published: (2025)
by: Lykov, Artem, et al.
Published: (2025)
FlightDiffusion: Revolutionising Autonomous Drone Training with Diffusion Models Generating FPV Video
by: Serpiva, Valerii, et al.
Published: (2025)
by: Serpiva, Valerii, et al.
Published: (2025)
DiffusionCinema: Text-to-Aerial Cinematography
by: Serpiva, Valerii, et al.
Published: (2026)
by: Serpiva, Valerii, et al.
Published: (2026)
EagleVision: A Multi-Task Benchmark for Cross-Domain Perception in High-Speed Autonomous Racing
by: Yagudin, Zakhar, et al.
Published: (2026)
by: Yagudin, Zakhar, et al.
Published: (2026)
HumanoidVLM: Vision-Language-Guided Impedance Control for Contact-Rich Humanoid Manipulation
by: Mahmoud, Yara, et al.
Published: (2026)
by: Mahmoud, Yara, et al.
Published: (2026)
LLM-MARS: Large Language Model for Behavior Tree Generation and NLP-enhanced Dialogue in Multi-Agent Robot Systems
by: Lykov, Artem, et al.
Published: (2023)
by: Lykov, Artem, et al.
Published: (2023)
Industry 6.0: New Generation of Industry driven by Generative AI and Swarm of Heterogeneous Robots
by: Lykov, Artem, et al.
Published: (2024)
by: Lykov, Artem, et al.
Published: (2024)
HapticVLM: VLM-Driven Texture Recognition Aimed at Intelligent Haptic Interaction
by: Khan, Muhammad Haris, et al.
Published: (2025)
by: Khan, Muhammad Haris, et al.
Published: (2025)
MissionGPT: Mission Planner for Mobile Robot based on Robotics Transformer Model
by: Berman, Vladimir, et al.
Published: (2024)
by: Berman, Vladimir, et al.
Published: (2024)
Collaborative Trajectory Prediction via Late Fusion
by: Madjid, Nadya Abdel, et al.
Published: (2026)
by: Madjid, Nadya Abdel, et al.
Published: (2026)
Evolution 6.0: Robot Evolution through Generative Design
by: Khan, Muhammad Haris, et al.
Published: (2025)
by: Khan, Muhammad Haris, et al.
Published: (2025)
GoalVLM: VLM-driven Object Goal Navigation for Multi-Agent System
by: James, MoniJesu, et al.
Published: (2026)
by: James, MoniJesu, et al.
Published: (2026)
VLM-UDMC: VLM-Enhanced Unified Decision-Making and Motion Control for Urban Autonomous Driving
by: Liu, Haichao, et al.
Published: (2025)
by: Liu, Haichao, et al.
Published: (2025)
DogSurf: Quadruped Robot Capable of GRU-based Surface Recognition for Blind Person Navigation
by: Bazhenov, Artem, et al.
Published: (2024)
by: Bazhenov, Artem, et al.
Published: (2024)
TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving
by: Li, Jiaxiang, et al.
Published: (2026)
by: Li, Jiaxiang, et al.
Published: (2026)
Bi-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Dexterous Manipulations
by: Gbagbe, Koffivi Fidèle, et al.
Published: (2024)
by: Gbagbe, Koffivi Fidèle, et al.
Published: (2024)
DreamToNav: Generalizable Navigation for Robots via Generative Video Planning
by: Serpiva, Valerii, et al.
Published: (2026)
by: Serpiva, Valerii, et al.
Published: (2026)
GenerativeMPC: VLM-RAG-guided Whole-Body MPC with Virtual Impedance for Bimanual Mobile Manipulation
by: Fernando, Marcelino Julio, et al.
Published: (2026)
by: Fernando, Marcelino Julio, et al.
Published: (2026)
UAV-VLRR: Vision-Language Informed NMPC for Rapid Response in UAV Search and Rescue
by: Yaqoot, Yasheerah, et al.
Published: (2025)
by: Yaqoot, Yasheerah, et al.
Published: (2025)
ImpedanceGPT: VLM-driven Impedance Control of Swarm of Mini-drones for Intelligent Navigation in Dynamic Environment
by: Batool, Faryal, et al.
Published: (2025)
by: Batool, Faryal, et al.
Published: (2025)
UAV-VLPA*: A Vision-Language-Path-Action System for Optimal Route Generation on a Large Scales
by: Sautenkov, Oleg, et al.
Published: (2025)
by: Sautenkov, Oleg, et al.
Published: (2025)
VLM-MPC: Vision Language Foundation Model (VLM)-Guided Model Predictive Controller (MPC) for Autonomous Driving
by: Long, Keke, et al.
Published: (2024)
by: Long, Keke, et al.
Published: (2024)
VERDI: VLM-Embedded Reasoning for Autonomous Driving
by: Feng, Bowen, et al.
Published: (2025)
by: Feng, Bowen, et al.
Published: (2025)
AnywhereVLA: Language-Conditioned Exploration and Mobile Manipulation
by: Gubernatorov, Konstantin, et al.
Published: (2025)
by: Gubernatorov, Konstantin, et al.
Published: (2025)
H-RINS: Hierarchical Tightly-coupled Radar-Inertial Navigation via Smoothing and Mapping
by: Abdulkarim, Ali Alridha, et al.
Published: (2026)
by: Abdulkarim, Ali Alridha, et al.
Published: (2026)
CognitiveDrone: A VLA Model and Evaluation Benchmark for Real-Time Cognitive Task Solving and Reasoning in UAVs
by: Lykov, Artem, et al.
Published: (2025)
by: Lykov, Artem, et al.
Published: (2025)
Similar Items
-
VDT-Auto: End-to-end Autonomous Driving with VLM-Guided Diffusion Transformers
by: Guo, Ziang, et al.
Published: (2025) -
FADet: A Multi-sensor 3D Object Detection Network based on Local Featured Attention
by: Guo, Ziang, et al.
Published: (2024) -
METDrive: Multi-modal End-to-end Autonomous Driving with Temporal Guidance
by: Guo, Ziang, et al.
Published: (2024) -
HawkDrive: A Transformer-driven Visual Perception System for Autonomous Driving in Night Scene
by: Guo, Ziang, et al.
Published: (2024) -
VR-GPT: Visual Language Model for Intelligent Virtual Reality Applications
by: Konenkov, Mikhail, et al.
Published: (2024)