VL-DPO: Vision-Language-Guided Finetuning for Preference-Aligned Autonomous Driving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Zhefan, Jerfel, Ghassen, Haliem, Marina, Zhao, Qi, Kang, Jeonhyung, Refaat, Khaled S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AVC-DPO: Aligned Video Captioning via Direct Preference Optimization
von: Tang, Jiyang, et al.
Veröffentlicht: (2025)
von: Tang, Jiyang, et al.
Veröffentlicht: (2025)
DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving
von: Shang, Shuyao, et al.
Veröffentlicht: (2025)
von: Shang, Shuyao, et al.
Veröffentlicht: (2025)
V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization
von: Xie, Yuxi, et al.
Veröffentlicht: (2024)
von: Xie, Yuxi, et al.
Veröffentlicht: (2024)
PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation
von: Huang, Qihan, et al.
Veröffentlicht: (2024)
von: Huang, Qihan, et al.
Veröffentlicht: (2024)
ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO
von: Ahn, Daechul, et al.
Veröffentlicht: (2024)
von: Ahn, Daechul, et al.
Veröffentlicht: (2024)
S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models
von: Shukla, Nitish, et al.
Veröffentlicht: (2026)
von: Shukla, Nitish, et al.
Veröffentlicht: (2026)
Driving with InternVL: Oustanding Champion in the Track on Driving with Language of the Autonomous Grand Challenge at CVPR 2024
von: Li, Jiahan, et al.
Veröffentlicht: (2024)
von: Li, Jiahan, et al.
Veröffentlicht: (2024)
CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs
von: Ouali, Yassine, et al.
Veröffentlicht: (2024)
von: Ouali, Yassine, et al.
Veröffentlicht: (2024)
When Preferences Diverge: Aligning Diffusion Models with Minority-Aware Adaptive DPO
von: Zhang, Lingfan, et al.
Veröffentlicht: (2025)
von: Zhang, Lingfan, et al.
Veröffentlicht: (2025)
DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
von: Tian, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Tian, Xiaoyu, et al.
Veröffentlicht: (2024)
GraSP-VL: Length as a Semantic Granularity Interface for Vision-Language Representations
von: Li, Zesheng, et al.
Veröffentlicht: (2026)
von: Li, Zesheng, et al.
Veröffentlicht: (2026)
Firebolt-VL: Efficient Vision-Language Understanding with Cross-Modality Modulation
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2026)
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2026)
VL-OrdinalFormer: Vision Language Guided Ordinal Transformers for Interpretable Knee Osteoarthritis Grading
von: Ullah, Zahid, et al.
Veröffentlicht: (2025)
von: Ullah, Zahid, et al.
Veröffentlicht: (2025)
DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving
von: Jiang, Anqing, et al.
Veröffentlicht: (2025)
von: Jiang, Anqing, et al.
Veröffentlicht: (2025)
SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model
von: Zhang, Yongting, et al.
Veröffentlicht: (2024)
von: Zhang, Yongting, et al.
Veröffentlicht: (2024)
VLP: Vision Language Planning for Autonomous Driving
von: Pan, Chenbin, et al.
Veröffentlicht: (2024)
von: Pan, Chenbin, et al.
Veröffentlicht: (2024)
MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models
von: Liu, Ziyu, et al.
Veröffentlicht: (2024)
von: Liu, Ziyu, et al.
Veröffentlicht: (2024)
DriveCritic: Towards Context-Aware, Human-Aligned Evaluation for Autonomous Driving with Vision-Language Models
von: Song, Jingyu, et al.
Veröffentlicht: (2025)
von: Song, Jingyu, et al.
Veröffentlicht: (2025)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
von: Chen, Zhe, et al.
Veröffentlicht: (2023)
von: Chen, Zhe, et al.
Veröffentlicht: (2023)
VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesis
von: Chu, Meng, et al.
Veröffentlicht: (2025)
von: Chu, Meng, et al.
Veröffentlicht: (2025)
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
von: Wu, Yanhao, et al.
Veröffentlicht: (2026)
von: Wu, Yanhao, et al.
Veröffentlicht: (2026)
Language-Aware Information Maximization for Transductive Few-Shot CLIP
von: Baklouti, Ghassen, et al.
Veröffentlicht: (2025)
von: Baklouti, Ghassen, et al.
Veröffentlicht: (2025)
Spatial-aware Vision Language Model for Autonomous Driving
von: Wei, Weijie, et al.
Veröffentlicht: (2025)
von: Wei, Weijie, et al.
Veröffentlicht: (2025)
Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
von: Xiong, Minhao, et al.
Veröffentlicht: (2025)
von: Xiong, Minhao, et al.
Veröffentlicht: (2025)
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
von: Cui, Erfei, et al.
Veröffentlicht: (2023)
von: Cui, Erfei, et al.
Veröffentlicht: (2023)
BREATH-VL: Vision-Language-Guided 6-DoF Bronchoscopy Localization via Semantic-Geometric Fusion
von: Tian, Qingyao, et al.
Veröffentlicht: (2026)
von: Tian, Qingyao, et al.
Veröffentlicht: (2026)
RealDPO: Real or Not Real, that is the Preference
von: Cheng, Guo, et al.
Veröffentlicht: (2025)
von: Cheng, Guo, et al.
Veröffentlicht: (2025)
Localization-Guided Foreground Augmentation in Autonomous Driving
von: Yong, Jiawei, et al.
Veröffentlicht: (2026)
von: Yong, Jiawei, et al.
Veröffentlicht: (2026)
PhysCorr: Dual-Reward DPO for Physics-Constrained Text-to-Video Generation with Automated Preference Selection
von: Wang, Peiyao, et al.
Veröffentlicht: (2025)
von: Wang, Peiyao, et al.
Veröffentlicht: (2025)
Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
von: Dong, Daxiang, et al.
Veröffentlicht: (2025)
von: Dong, Daxiang, et al.
Veröffentlicht: (2025)
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
von: Zhao, Zongchuang, et al.
Veröffentlicht: (2025)
von: Zhao, Zongchuang, et al.
Veröffentlicht: (2025)
Thermo-VL: Extending Vision-Language Models to Thermal Infrared Perception
von: Thushara, Rusiru, et al.
Veröffentlicht: (2026)
von: Thushara, Rusiru, et al.
Veröffentlicht: (2026)
VL-Reader: Vision and Language Reconstructor is an Effective Scene Text Recognizer
von: Zhong, Humen, et al.
Veröffentlicht: (2024)
von: Zhong, Humen, et al.
Veröffentlicht: (2024)
3VL: Using Trees to Improve Vision-Language Models' Interpretability
von: Yellinek, Nir, et al.
Veröffentlicht: (2023)
von: Yellinek, Nir, et al.
Veröffentlicht: (2023)
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
von: Wang, Shijing, et al.
Veröffentlicht: (2025)
von: Wang, Shijing, et al.
Veröffentlicht: (2025)
ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning
von: Hao, Zhiwei, et al.
Veröffentlicht: (2024)
von: Hao, Zhiwei, et al.
Veröffentlicht: (2024)
Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key
von: Yang, Zhihe, et al.
Veröffentlicht: (2025)
von: Yang, Zhihe, et al.
Veröffentlicht: (2025)
LaCoVL-FER: Landmark-Guided Contrastive Learning Network with Vision-Language Enhancement for Facial Expression Recognition
von: Wang, Jiaxin, et al.
Veröffentlicht: (2026)
von: Wang, Jiaxin, et al.
Veröffentlicht: (2026)
Visual Adversarial Attack on Vision-Language Models for Autonomous Driving
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2024)
VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
von: Seong, Hyunki, et al.
Veröffentlicht: (2025)
von: Seong, Hyunki, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AVC-DPO: Aligned Video Captioning via Direct Preference Optimization
von: Tang, Jiyang, et al.
Veröffentlicht: (2025) -
DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving
von: Shang, Shuyao, et al.
Veröffentlicht: (2025) -
V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization
von: Xie, Yuxi, et al.
Veröffentlicht: (2024) -
PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation
von: Huang, Qihan, et al.
Veröffentlicht: (2024) -
ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO
von: Ahn, Daechul, et al.
Veröffentlicht: (2024)