LLaViDA: A Large Language Vision Driving Assistant for Explicit Reasoning and Enhanced Trajectory Planning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yudong, Hallyburton, Spencer, Kim, Jiwoo, Lin, Yueqian, Li, Yiming, Wang, Qinsi, Ye, Hui, Sun, Jingwei, Pajic, Miroslav, Chen, Yiran, Li, Hai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bayesian Methods for Trust in Collaborative Multi-Agent Autonomy
by: Hallyburton, R. Spencer, et al.
Published: (2024)
by: Hallyburton, R. Spencer, et al.
Published: (2024)
Assured Autonomy with Neuro-Symbolic Perception
by: Hallyburton, R. Spencer, et al.
Published: (2025)
by: Hallyburton, R. Spencer, et al.
Published: (2025)
Security-Aware Sensor Fusion with MATE: the Multi-Agent Trust Estimator
by: Hallyburton, R. Spencer, et al.
Published: (2025)
by: Hallyburton, R. Spencer, et al.
Published: (2025)
Trusted Data Fusion, Multi-Agent Autonomy, Autonomous Vehicles
by: Hallyburton, R. Spencer, et al.
Published: (2025)
by: Hallyburton, R. Spencer, et al.
Published: (2025)
Keyframe-oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-Form Video Processing
by: Liu, Yudong, et al.
Published: (2025)
by: Liu, Yudong, et al.
Published: (2025)
A Multi-Agent Security Testbed for the Analysis of Attacks and Defenses in Collaborative Sensor Fusion
by: Hallyburton, R. Spencer, et al.
Published: (2024)
by: Hallyburton, R. Spencer, et al.
Published: (2024)
AsyncVoice Agent: Real-Time Explanation for LLM Planning and Reasoning
by: Lin, Yueqian, et al.
Published: (2025)
by: Lin, Yueqian, et al.
Published: (2025)
Probabilistic Segmentation for Robust Field of View Estimation
by: Hallyburton, R. Spencer, et al.
Published: (2025)
by: Hallyburton, R. Spencer, et al.
Published: (2025)
What Would Trojans Do? Exploiting Partial-Information Vulnerabilities in Autonomous Vehicle Sensing
by: Hallyburton, R. Spencer, et al.
Published: (2023)
by: Hallyburton, R. Spencer, et al.
Published: (2023)
HippoMM: Hippocampal-inspired Multimodal Memory for Long Audiovisual Event Understanding
by: Lin, Yueqian, et al.
Published: (2025)
by: Lin, Yueqian, et al.
Published: (2025)
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
by: Lin, Yueqian, et al.
Published: (2025)
by: Lin, Yueqian, et al.
Published: (2025)
RaGNNarok: A Light-Weight Graph Neural Network for Enhancing Radar Point Clouds on Unmanned Ground Vehicles
by: Hunt, David, et al.
Published: (2025)
by: Hunt, David, et al.
Published: (2025)
Semantic Area Graph Reasoning for Multi-Robot Language-Guided Search
by: Wang, Ruiyang, et al.
Published: (2026)
by: Wang, Ruiyang, et al.
Published: (2026)
RadCloud: Real-Time High-Resolution Point Cloud Generation Using Low-Cost Radars for Aerial and Ground Vehicles
by: Hunt, David, et al.
Published: (2024)
by: Hunt, David, et al.
Published: (2024)
CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
NNiT: Width-Agnostic Neural Network Generation with Structurally Aligned Weight Spaces
by: Kim, Jiwoo, et al.
Published: (2026)
by: Kim, Jiwoo, et al.
Published: (2026)
Latent Bridge: Feature Delta Prediction for Efficient Dual-System Vision-Language-Action Model Inference
by: Liu, Yudong, et al.
Published: (2026)
by: Liu, Yudong, et al.
Published: (2026)
Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
COMRES-VLM: Coordinated Multi-Robot Exploration and Search using Vision Language Models
by: Wang, Ruiyang, et al.
Published: (2025)
by: Wang, Ruiyang, et al.
Published: (2025)
FlashFPS: Efficient Farthest Point Sampling for Large-Scale Point Clouds via Pruning and Caching
by: Fu, Yuzhe, et al.
Published: (2026)
by: Fu, Yuzhe, et al.
Published: (2026)
SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval
by: Lin, Yueqian, et al.
Published: (2024)
by: Lin, Yueqian, et al.
Published: (2024)
SD-NAE: Generating Natural Adversarial Examples with Stable Diffusion
by: Lin, Yueqian, et al.
Published: (2023)
by: Lin, Yueqian, et al.
Published: (2023)
Secure Planning Against Stealthy Attacks via Model-Free Reinforcement Learning
by: Bozkurt, Alper Kamil, et al.
Published: (2020)
by: Bozkurt, Alper Kamil, et al.
Published: (2020)
Yo'LLaVA: Your Personalized Language and Vision Assistant
by: Nguyen, Thao, et al.
Published: (2024)
by: Nguyen, Thao, et al.
Published: (2024)
ReGraP-LLaVA: Reasoning enabled Graph-based Personalized Large Language and Vision Assistant
by: Xiang, Yifan, et al.
Published: (2025)
by: Xiang, Yifan, et al.
Published: (2025)
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
by: Wang, Yiheng, et al.
Published: (2026)
by: Wang, Yiheng, et al.
Published: (2026)
Bridging the Perception Gap: A Lightweight Coarse-to-Fine Architecture for Edge Audio Systems
by: Zhang, Hengfan, et al.
Published: (2026)
by: Zhang, Hengfan, et al.
Published: (2026)
LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound
by: Guo, Xuechen, et al.
Published: (2024)
by: Guo, Xuechen, et al.
Published: (2024)
Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
by: Wei, Chiyue, et al.
Published: (2025)
by: Wei, Chiyue, et al.
Published: (2025)
Angles Don't Lie: Unlocking Training-Efficient RL Through the Model's Own Signals
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
by: Lu, Weiheng, et al.
Published: (2024)
by: Lu, Weiheng, et al.
Published: (2024)
ViLLa: Video Reasoning Segmentation with Large Language Model
by: Zheng, Rongkun, et al.
Published: (2024)
by: Zheng, Rongkun, et al.
Published: (2024)
Black-box Stealthy GPS Attacks on Unmanned Aerial Vehicles
by: Khazraei, Amir, et al.
Published: (2024)
by: Khazraei, Amir, et al.
Published: (2024)
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
by: Jin, Yizhang, et al.
Published: (2024)
by: Jin, Yizhang, et al.
Published: (2024)
MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models
by: Wang, Zhongxi, et al.
Published: (2026)
by: Wang, Zhongxi, et al.
Published: (2026)
Power-LLaVA: Large Language and Vision Assistant for Power Transmission Line Inspection
by: Wang, Jiahao, et al.
Published: (2024)
by: Wang, Jiahao, et al.
Published: (2024)
T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning
by: Wang, Qinsi, et al.
Published: (2026)
by: Wang, Qinsi, et al.
Published: (2026)
LLaRA: Large Language-Recommendation Assistant
by: Liao, Jiayi, et al.
Published: (2023)
by: Liao, Jiayi, et al.
Published: (2023)
SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant
by: Sun, Guohao, et al.
Published: (2024)
by: Sun, Guohao, et al.
Published: (2024)
LLaVA-LE: Large Language-and-Vision Assistant for Lunar Exploration
by: Inal, Gokce, et al.
Published: (2026)
by: Inal, Gokce, et al.
Published: (2026)
Similar Items
-
Bayesian Methods for Trust in Collaborative Multi-Agent Autonomy
by: Hallyburton, R. Spencer, et al.
Published: (2024) -
Assured Autonomy with Neuro-Symbolic Perception
by: Hallyburton, R. Spencer, et al.
Published: (2025) -
Security-Aware Sensor Fusion with MATE: the Multi-Agent Trust Estimator
by: Hallyburton, R. Spencer, et al.
Published: (2025) -
Trusted Data Fusion, Multi-Agent Autonomy, Autonomous Vehicles
by: Hallyburton, R. Spencer, et al.
Published: (2025) -
Keyframe-oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-Form Video Processing
by: Liu, Yudong, et al.
Published: (2025)