Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Hao, Gao, Zhanning, Chen, Zhili, Ye, Maosheng, Chen, Qifeng, Cao, Tongyi, Qi, Honggang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PPAD: Iterative Interactions of Prediction and Planning for End-to-end Autonomous Driving
by: Chen, Zhili, et al.
Published: (2023)
by: Chen, Zhili, et al.
Published: (2023)
Cross-Cluster Shifting for Efficient and Effective 3D Object Detection in Autonomous Driving
by: Chen, Zhili, et al.
Published: (2024)
by: Chen, Zhili, et al.
Published: (2024)
Learning High-resolution Vector Representation from Multi-Camera Images for 3D Object Detection
by: Chen, Zhili, et al.
Published: (2024)
by: Chen, Zhili, et al.
Published: (2024)
Hint-AD: Holistically Aligned Interpretability in End-to-End Autonomous Driving
by: Ding, Kairui, et al.
Published: (2024)
by: Ding, Kairui, et al.
Published: (2024)
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
by: Mei, Xiaodong, et al.
Published: (2026)
by: Mei, Xiaodong, et al.
Published: (2026)
VGGT-MPR: VGGT-Enhanced Multimodal Place Recognition in Autonomous Driving Environments
by: Xu, Jingyi, et al.
Published: (2026)
by: Xu, Jingyi, et al.
Published: (2026)
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
by: Li, Honglin, et al.
Published: (2024)
by: Li, Honglin, et al.
Published: (2024)
Enhancing Multimodal Large Language Models with Multi-instance Visual Prompt Generator for Visual Representation Enrichment
by: Zhong, Wenliang, et al.
Published: (2024)
by: Zhong, Wenliang, et al.
Published: (2024)
ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
by: Zhang, Yuan, et al.
Published: (2025)
by: Zhang, Yuan, et al.
Published: (2025)
Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving
by: Zhou, Yuhan, et al.
Published: (2026)
by: Zhou, Yuhan, et al.
Published: (2026)
Video Token Sparsification for Efficient Multimodal LLMs in Autonomous Driving
by: Ma, Yunsheng, et al.
Published: (2024)
by: Ma, Yunsheng, et al.
Published: (2024)
CoT-Drive: Efficient Motion Forecasting for Autonomous Driving with LLMs and Chain-of-Thought Prompting
by: Liao, Haicheng, et al.
Published: (2025)
by: Liao, Haicheng, et al.
Published: (2025)
ControlLoc: Physical-World Hijacking Attack on Visual Perception in Autonomous Driving
by: Ma, Chen, et al.
Published: (2024)
by: Ma, Chen, et al.
Published: (2024)
Efficient Visual Question Answering Pipeline for Autonomous Driving via Scene Region Compression
by: Cai, Yuliang, et al.
Published: (2026)
by: Cai, Yuliang, et al.
Published: (2026)
Language Prompt for Autonomous Driving
by: Wu, Dongming, et al.
Published: (2023)
by: Wu, Dongming, et al.
Published: (2023)
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
by: Zhou, Yikang, et al.
Published: (2025)
by: Zhou, Yikang, et al.
Published: (2025)
SPIRE: Semantic Prompt-Driven Image Restoration
by: Qi, Chenyang, et al.
Published: (2023)
by: Qi, Chenyang, et al.
Published: (2023)
InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving
by: Song, Ruiqi, et al.
Published: (2025)
by: Song, Ruiqi, et al.
Published: (2025)
Toward Unified Multimodal Representation Learning for Autonomous Driving
by: Tao, Ximeng, et al.
Published: (2026)
by: Tao, Ximeng, et al.
Published: (2026)
Visual Prompting in LLMs for Enhancing Emotion Recognition
by: Zhang, Qixuan, et al.
Published: (2024)
by: Zhang, Qixuan, et al.
Published: (2024)
LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models
by: Tian, Shi-Yu, et al.
Published: (2026)
by: Tian, Shi-Yu, et al.
Published: (2026)
Retrieval-Enhanced Visual Prompt Learning for Few-shot Classification
by: Rong, Jintao, et al.
Published: (2023)
by: Rong, Jintao, et al.
Published: (2023)
WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model
by: Zhang, Songyan, et al.
Published: (2024)
by: Zhang, Songyan, et al.
Published: (2024)
WHALES: A Multi-Agent Scheduling Dataset for Enhanced Cooperation in Autonomous Driving
by: Wang, Yinsong, et al.
Published: (2024)
by: Wang, Yinsong, et al.
Published: (2024)
The Devil is in the Few Shots: Iterative Visual Knowledge Completion for Few-shot Learning
by: Li, Yaohui, et al.
Published: (2024)
by: Li, Yaohui, et al.
Published: (2024)
Multimodal-Enhanced Objectness Learner for Corner Case Detection in Autonomous Driving
by: Xiao, Lixing, et al.
Published: (2024)
by: Xiao, Lixing, et al.
Published: (2024)
GraSP-VL: Length as a Semantic Granularity Interface for Vision-Language Representations
by: Li, Zesheng, et al.
Published: (2026)
by: Li, Zesheng, et al.
Published: (2026)
Multi-dimensional Visual Prompt Enhanced Image Restoration via Mamba-Transformer Aggregation
by: Jiang, Aiwen, et al.
Published: (2024)
by: Jiang, Aiwen, et al.
Published: (2024)
GSPR: Multimodal Place Recognition Using 3D Gaussian Splatting for Autonomous Driving
by: Qi, Zhangshuo, et al.
Published: (2024)
by: Qi, Zhangshuo, et al.
Published: (2024)
SegLocNet: Multimodal Localization Network for Autonomous Driving via Bird's-Eye-View Segmentation
by: Zhou, Zijie, et al.
Published: (2025)
by: Zhou, Zijie, et al.
Published: (2025)
Visual Point Cloud Forecasting enables Scalable Autonomous Driving
by: Yang, Zetong, et al.
Published: (2023)
by: Yang, Zetong, et al.
Published: (2023)
SlowPerception: Physical-World Latency Attack against Visual Perception in Autonomous Driving
by: Ma, Chen, et al.
Published: (2024)
by: Ma, Chen, et al.
Published: (2024)
SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation
by: Wang, Ruoyu, et al.
Published: (2026)
by: Wang, Ruoyu, et al.
Published: (2026)
Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2024)
by: Jia, Xiaosong, et al.
Published: (2024)
Diffusion-Based Generative Models for 3D Occupancy Prediction in Autonomous Driving
by: Wang, Yunshen, et al.
Published: (2025)
by: Wang, Yunshen, et al.
Published: (2025)
Enhancing Prompt Following with Visual Control Through Training-Free Mask-Guided Diffusion
by: Chen, Hongyu, et al.
Published: (2024)
by: Chen, Hongyu, et al.
Published: (2024)
Customized Visual Storytelling with Unified Multimodal LLMs
by: Li, Wei-Hua, et al.
Published: (2026)
by: Li, Wei-Hua, et al.
Published: (2026)
DriveVGGT: Calibration-Constrained Visual Geometry Transformers for Multi-Camera Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2025)
by: Jia, Xiaosong, et al.
Published: (2025)
Enhancing Medical Visual Grounding via Knowledge-guided Spatial Prompts
by: Gao, Yifan, et al.
Published: (2026)
by: Gao, Yifan, et al.
Published: (2026)
Lightning NeRF: Efficient Hybrid Scene Representation for Autonomous Driving
by: Cao, Junyi, et al.
Published: (2024)
by: Cao, Junyi, et al.
Published: (2024)
Similar Items
-
PPAD: Iterative Interactions of Prediction and Planning for End-to-end Autonomous Driving
by: Chen, Zhili, et al.
Published: (2023) -
Cross-Cluster Shifting for Efficient and Effective 3D Object Detection in Autonomous Driving
by: Chen, Zhili, et al.
Published: (2024) -
Learning High-resolution Vector Representation from Multi-Camera Images for 3D Object Detection
by: Chen, Zhili, et al.
Published: (2024) -
Hint-AD: Holistically Aligned Interpretability in End-to-End Autonomous Driving
by: Ding, Kairui, et al.
Published: (2024) -
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
by: Mei, Xiaodong, et al.
Published: (2026)