CarLLaVA: Vision language models for camera-only closed-loop driving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Renz, Katrin, Chen, Long, Marcu, Ana-Maria, Hünermann, Jan, Hanotte, Benoit, Karnsund, Alice, Shotton, Jamie, Arani, Elahe, Sinavski, Oleg |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LingoQA: Visual Question Answering for Autonomous Driving
von: Marcu, Ana-Maria, et al.
Veröffentlicht: (2023)
von: Marcu, Ana-Maria, et al.
Veröffentlicht: (2023)
SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment
von: Renz, Katrin, et al.
Veröffentlicht: (2025)
von: Renz, Katrin, et al.
Veröffentlicht: (2025)
GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
von: Russell, Lloyd, et al.
Veröffentlicht: (2025)
von: Russell, Lloyd, et al.
Veröffentlicht: (2025)
El Vaticano II, ¿texto constitucional de la fe? Una carta de Peter Hünermann
von: Peter Hünermann
Veröffentlicht: (2016)
von: Peter Hünermann
Veröffentlicht: (2016)
Robotic Learning in your Backyard: A Neural Simulator from Open Source Components
von: Zhou, Liyou, et al.
Veröffentlicht: (2024)
von: Zhou, Liyou, et al.
Veröffentlicht: (2024)
EfficientLLaVA:Generalizable Auto-Pruning for Large Vision-language Models
von: Liang, Yinan, et al.
Veröffentlicht: (2025)
von: Liang, Yinan, et al.
Veröffentlicht: (2025)
NOVO: Bridging LLaVA and SAM with Visual-only Prompts for Reasoning Segmentation
von: Yoon, Kyung-Yoon, et al.
Veröffentlicht: (2025)
von: Yoon, Kyung-Yoon, et al.
Veröffentlicht: (2025)
Conserve-Update-Revise to Cure Generalization and Robustness Trade-off in Adversarial Training
von: Gowda, Shruthi, et al.
Veröffentlicht: (2024)
von: Gowda, Shruthi, et al.
Veröffentlicht: (2024)
Gradual Divergence for Seamless Adaptation: A Novel Domain Incremental Learning Method
von: Jeeveswaran, Kishaan, et al.
Veröffentlicht: (2024)
von: Jeeveswaran, Kishaan, et al.
Veröffentlicht: (2024)
Beyond Unimodal Learning: The Importance of Integrating Multiple Modalities for Lifelong Learning
von: Sarfraz, Fahad, et al.
Veröffentlicht: (2024)
von: Sarfraz, Fahad, et al.
Veröffentlicht: (2024)
Can We Break Free from Strong Data Augmentations in Self-Supervised Learning?
von: Gowda, Shruthi, et al.
Veröffentlicht: (2024)
von: Gowda, Shruthi, et al.
Veröffentlicht: (2024)
AstroLLaVA: towards the unification of astronomical data and natural language
von: Zaman, Sharaf, et al.
Veröffentlicht: (2025)
von: Zaman, Sharaf, et al.
Veröffentlicht: (2025)
Untangling complex ethical issues involving wildlife
von: Justine Shotton
Veröffentlicht: (2024)
von: Justine Shotton
Veröffentlicht: (2024)
Welfare and wildlife rehabilitation
von: Justine Shotton
Veröffentlicht: (2026)
von: Justine Shotton
Veröffentlicht: (2026)
Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification
von: Huang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Huang, Wenxuan, et al.
Veröffentlicht: (2024)
The Effectiveness of Random Forgetting for Robust Generalization
von: Ramkumar, Vijaya Raghavan T, et al.
Veröffentlicht: (2024)
von: Ramkumar, Vijaya Raghavan T, et al.
Veröffentlicht: (2024)
Can Sound Replace Vision in LLaVA With Token Substitution?
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
Yo'LLaVA: Your Personalized Language and Vision Assistant
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
LLaVA-OneVision: Easy Visual Task Transfer
von: Li, Bo, et al.
Veröffentlicht: (2024)
von: Li, Bo, et al.
Veröffentlicht: (2024)
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
von: Zeer, Ahmed, et al.
Veröffentlicht: (2024)
von: Zeer, Ahmed, et al.
Veröffentlicht: (2024)
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
von: Yan, Dawei, et al.
Veröffentlicht: (2024)
von: Yan, Dawei, et al.
Veröffentlicht: (2024)
SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant
von: Sun, Guohao, et al.
Veröffentlicht: (2024)
von: Sun, Guohao, et al.
Veröffentlicht: (2024)
LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound
von: Guo, Xuechen, et al.
Veröffentlicht: (2024)
von: Guo, Xuechen, et al.
Veröffentlicht: (2024)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
von: An, Ruichuan, et al.
Veröffentlicht: (2024)
von: An, Ruichuan, et al.
Veröffentlicht: (2024)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
von: An, Ruichuan, et al.
Veröffentlicht: (2025)
von: An, Ruichuan, et al.
Veröffentlicht: (2025)
LLaVA-LE: Large Language-and-Vision Assistant for Lunar Exploration
von: Inal, Gokce, et al.
Veröffentlicht: (2026)
von: Inal, Gokce, et al.
Veröffentlicht: (2026)
X-LLaVA: Optimizing Bilingual Large Vision-Language Alignment
von: Shin, Dongjae, et al.
Veröffentlicht: (2024)
von: Shin, Dongjae, et al.
Veröffentlicht: (2024)
IMEX-Reg: Implicit-Explicit Regularization in the Function Space for Continual Learning
von: Bhat, Prashant, et al.
Veröffentlicht: (2024)
von: Bhat, Prashant, et al.
Veröffentlicht: (2024)
Mitigating Interference in the Knowledge Continuum through Attention-Guided Incremental Learning
von: Bhat, Prashant, et al.
Veröffentlicht: (2024)
von: Bhat, Prashant, et al.
Veröffentlicht: (2024)
Continual Learning Beyond Experience Rehearsal and Full Model Surrogates
von: Bhat, Prashant, et al.
Veröffentlicht: (2025)
von: Bhat, Prashant, et al.
Veröffentlicht: (2025)
PhysVid: Physics Aware Local Conditioning for Generative Video Models
von: Pathak, Saurabh, et al.
Veröffentlicht: (2026)
von: Pathak, Saurabh, et al.
Veröffentlicht: (2026)
LangProp: A code optimization framework using Large Language Models applied to driving
von: Ishida, Shu, et al.
Veröffentlicht: (2024)
von: Ishida, Shu, et al.
Veröffentlicht: (2024)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
von: Shu, Fangxun, et al.
Veröffentlicht: (2024)
von: Shu, Fangxun, et al.
Veröffentlicht: (2024)
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
von: Jahagirdar, Soumya, et al.
Veröffentlicht: (2026)
von: Jahagirdar, Soumya, et al.
Veröffentlicht: (2026)
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
von: Jin, Yizhang, et al.
Veröffentlicht: (2024)
von: Jin, Yizhang, et al.
Veröffentlicht: (2024)
Space-LLaVA: a Vision-Language Model Adapted to Extraterrestrial Applications
von: Foutter, Matthew, et al.
Veröffentlicht: (2024)
von: Foutter, Matthew, et al.
Veröffentlicht: (2024)
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
von: Lu, Weiheng, et al.
Veröffentlicht: (2024)
von: Lu, Weiheng, et al.
Veröffentlicht: (2024)
ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
von: Ye, Xubing, et al.
Veröffentlicht: (2024)
von: Ye, Xubing, et al.
Veröffentlicht: (2024)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
von: Lin, Bin, et al.
Veröffentlicht: (2024)
von: Lin, Bin, et al.
Veröffentlicht: (2024)
Why do LLaVA Vision-Language Models Reply to Images in English?
von: Hinck, Musashi, et al.
Veröffentlicht: (2024)
von: Hinck, Musashi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LingoQA: Visual Question Answering for Autonomous Driving
von: Marcu, Ana-Maria, et al.
Veröffentlicht: (2023) -
SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment
von: Renz, Katrin, et al.
Veröffentlicht: (2025) -
GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
von: Russell, Lloyd, et al.
Veröffentlicht: (2025) -
El Vaticano II, ¿texto constitucional de la fe? Una carta de Peter Hünermann
von: Peter Hünermann
Veröffentlicht: (2016) -
Robotic Learning in your Backyard: A Neural Simulator from Open Source Components
von: Zhou, Liyou, et al.
Veröffentlicht: (2024)