VLM-VPI: A Vision-Language Reasoning Framework for Improving Automated Vehicle-Pedestrian Interactions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pu, Qingwen, Xie, Kun, Liu, Yuxiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026)
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
von: Liu, Bingnan, et al.
Veröffentlicht: (2026)
von: Liu, Bingnan, et al.
Veröffentlicht: (2026)
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
von: Wu, Jason, et al.
Veröffentlicht: (2026)
von: Wu, Jason, et al.
Veröffentlicht: (2026)
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
von: Hou, Zhangcheng, et al.
Veröffentlicht: (2026)
von: Hou, Zhangcheng, et al.
Veröffentlicht: (2026)
Smooth regularization for efficient video recognition
von: Goldman, Gil, et al.
Veröffentlicht: (2025)
von: Goldman, Gil, et al.
Veröffentlicht: (2025)
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
von: Bergkvist, Viktor, et al.
Veröffentlicht: (2026)
von: Bergkvist, Viktor, et al.
Veröffentlicht: (2026)
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
von: Ji, Binbin, et al.
Veröffentlicht: (2025)
von: Ji, Binbin, et al.
Veröffentlicht: (2025)
Perceptual Flow Network for Visually Grounded Reasoning
von: Li, Yangfu, et al.
Veröffentlicht: (2026)
von: Li, Yangfu, et al.
Veröffentlicht: (2026)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
Survey Transfer Learning: Recycling Data with Silicon Responses
von: Amini, Ali
Veröffentlicht: (2025)
von: Amini, Ali
Veröffentlicht: (2025)
Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
von: Arnaud, Sergio, et al.
Veröffentlicht: (2025)
von: Arnaud, Sergio, et al.
Veröffentlicht: (2025)
FeedbackSTS-Det: Sparse Frames-Based Spatio-Temporal Semantic Feedback Network for Moving Infrared Small Target Detection
von: Huang, Yian, et al.
Veröffentlicht: (2026)
von: Huang, Yian, et al.
Veröffentlicht: (2026)
SERA-H: Beyond Native Sentinel Spatial Limits for High-Resolution Canopy Height Mapping
von: Boudras, Thomas, et al.
Veröffentlicht: (2025)
von: Boudras, Thomas, et al.
Veröffentlicht: (2025)
MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation
von: Bartkowiak, Patryk, et al.
Veröffentlicht: (2026)
von: Bartkowiak, Patryk, et al.
Veröffentlicht: (2026)
Implementing Adaptations for Vision AutoRegressive Model
von: Shaikh, Kaif, et al.
Veröffentlicht: (2025)
von: Shaikh, Kaif, et al.
Veröffentlicht: (2025)
Caption-Driven Explainability: Probing CNNs for Bias via CLIP
von: Koller, Patrick, et al.
Veröffentlicht: (2025)
von: Koller, Patrick, et al.
Veröffentlicht: (2025)
OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models
von: Koroglu, Mathis, et al.
Veröffentlicht: (2024)
von: Koroglu, Mathis, et al.
Veröffentlicht: (2024)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
von: Yang, Shan
Veröffentlicht: (2026)
von: Yang, Shan
Veröffentlicht: (2026)
In Context Learning with Vision Transformers: Case Study
von: Zhao, Antony, et al.
Veröffentlicht: (2025)
von: Zhao, Antony, et al.
Veröffentlicht: (2025)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
Automated Pollen Recognition in Optical and Holographic Microscopy Images
von: Warshaneyan, Swarn Singh, et al.
Veröffentlicht: (2025)
von: Warshaneyan, Swarn Singh, et al.
Veröffentlicht: (2025)
Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
von: Lin, Xiaojian, et al.
Veröffentlicht: (2025)
von: Lin, Xiaojian, et al.
Veröffentlicht: (2025)
From eye to AI: studying rodent social behavior in the era of machine Learning
von: Chindemi, Giuseppe, et al.
Veröffentlicht: (2025)
von: Chindemi, Giuseppe, et al.
Veröffentlicht: (2025)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
von: Chen, Yuangong, et al.
Veröffentlicht: (2026)
von: Chen, Yuangong, et al.
Veröffentlicht: (2026)
Object detection in adverse weather conditions for autonomous vehicles using Instruct Pix2Pix
von: Gurbindo, Unai, et al.
Veröffentlicht: (2025)
von: Gurbindo, Unai, et al.
Veröffentlicht: (2025)
Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing
von: Li, Zhuowei, et al.
Veröffentlicht: (2025)
von: Li, Zhuowei, et al.
Veröffentlicht: (2025)
SpectralCA: Bi-Directional Cross-Attention for Next-Generation UAV Hyperspectral Vision
von: Brovko, D. V.
Veröffentlicht: (2025)
von: Brovko, D. V.
Veröffentlicht: (2025)
Leum-VL Technical Report
von: He, Yuxuan, et al.
Veröffentlicht: (2026)
von: He, Yuxuan, et al.
Veröffentlicht: (2026)
Learning the meanings of function words from grounded language using a visual question answering model
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
Self-Supervised Polyp Re-Identification in Colonoscopy
von: Intrator, Yotam, et al.
Veröffentlicht: (2023)
von: Intrator, Yotam, et al.
Veröffentlicht: (2023)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
von: Gupta, Sunny, et al.
Veröffentlicht: (2024)
von: Gupta, Sunny, et al.
Veröffentlicht: (2024)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
von: Kashyap, Pankhi, et al.
Veröffentlicht: (2024)
von: Kashyap, Pankhi, et al.
Veröffentlicht: (2024)
QoSGMAA: A Robust Multi-Order Graph Attention and Adversarial Framework for Sparse QoS Prediction
von: Du, Guanchen, et al.
Veröffentlicht: (2025)
von: Du, Guanchen, et al.
Veröffentlicht: (2025)
DRIFT open dataset: A drone-derived intelligence for traffic analysis in urban environment
von: Lee, Hyejin, et al.
Veröffentlicht: (2025)
von: Lee, Hyejin, et al.
Veröffentlicht: (2025)
Modeling Vehicle-Type-Specific Pedestrian Crash Avoidance Behavior in Safety-Critical Interactions Using Smooth-Mamba Deep Reinforcement Learning
von: Pu, Qingwen, et al.
Veröffentlicht: (2026)
von: Pu, Qingwen, et al.
Veröffentlicht: (2026)
Efficient Attention: Attention with Linear Complexities
von: Shen, Zhuoran, et al.
Veröffentlicht: (2018)
von: Shen, Zhuoran, et al.
Veröffentlicht: (2018)
METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark
von: Yang, Xu, et al.
Veröffentlicht: (2025)
von: Yang, Xu, et al.
Veröffentlicht: (2025)
HATL: Hierarchical Adaptive-Transfer Learning Framework for Sign Language Machine Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2026)
von: Shahin, Nada, et al.
Veröffentlicht: (2026)
An Analysis of Layer-Freezing Strategies for Enhanced Transfer Learning in YOLO Architectures
von: Dobrzycki, Andrzej D., et al.
Veröffentlicht: (2025)
von: Dobrzycki, Andrzej D., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026) -
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025) -
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
von: Liu, Bingnan, et al.
Veröffentlicht: (2026) -
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
von: Wu, Jason, et al.
Veröffentlicht: (2026) -
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
von: Hou, Zhangcheng, et al.
Veröffentlicht: (2026)