Visual Reasoning at Urban Intersections: FineTuning GPT-4o for Traffic Conflict Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Masri, Sari, Ashqar, Huthaifa I., Elhenawy, Mohammed |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Large Language Models (LLMs) for Traffic Management at Urban Intersections: The Case of Mixed Traffic Scenarios
von: Masri, Sari, et al.
Veröffentlicht: (2024)
von: Masri, Sari, et al.
Veröffentlicht: (2024)
Large Language Models (LLMs) as Traffic Control Systems at Urban Intersections: A New Paradigm
von: Masri, Sari, et al.
Veröffentlicht: (2024)
von: Masri, Sari, et al.
Veröffentlicht: (2024)
HazardNet: A Small-Scale Vision Language Model for Real-Time Traffic Safety Detection at Edge Devices
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2025)
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2025)
Multimodal Large Language Models for Enhanced Traffic Safety: A Comprehensive Review and Future Trends
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2025)
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2025)
Using Multimodal Large Language Models for Automated Detection of Traffic Safety Critical Events
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2024)
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2024)
Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing
von: Ashqar, Huthaifa I., et al.
Veröffentlicht: (2024)
von: Ashqar, Huthaifa I., et al.
Veröffentlicht: (2024)
The Use of Multimodal Large Language Models to Detect Objects from Thermal Images: Transportation Applications
von: Ashqar, Huthaifa I., et al.
Veröffentlicht: (2024)
von: Ashqar, Huthaifa I., et al.
Veröffentlicht: (2024)
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
Object Detection using Oriented Window Learning Vi-sion Transformer: Roadway Assets Recognition
von: Alhadidi, Taqwa, et al.
Veröffentlicht: (2024)
von: Alhadidi, Taqwa, et al.
Veröffentlicht: (2024)
Transformer Models in Education: Summarizing Science Textbooks with AraBART, MT5, AraT5, and mBART
von: Masri, Sari, et al.
Veröffentlicht: (2024)
von: Masri, Sari, et al.
Veröffentlicht: (2024)
Enhancing Pavement Crack Classification with Bidirectional Cascaded Neural Networks
von: Alhadidi, Taqwa I., et al.
Veröffentlicht: (2025)
von: Alhadidi, Taqwa I., et al.
Veröffentlicht: (2025)
Automated Question Generation for Science Tests in Arabic Language Using NLP Techniques
von: Tami, Mohammad, et al.
Veröffentlicht: (2024)
von: Tami, Mohammad, et al.
Veröffentlicht: (2024)
Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
Automated Pavement Cracks Detection and Classification Using Deep Learning
von: Nafaa, Selvia, et al.
Veröffentlicht: (2024)
von: Nafaa, Selvia, et al.
Veröffentlicht: (2024)
Advancing Roadway Sign Detection with YOLO Models and Transfer Learning
von: Nafaa, Selvia, et al.
Veröffentlicht: (2024)
von: Nafaa, Selvia, et al.
Veröffentlicht: (2024)
Question-Answering (QA) Model for a Personalized Learning Assistant for Arabic Language
von: Sammoudi, Mohammad, et al.
Veröffentlicht: (2024)
von: Sammoudi, Mohammad, et al.
Veröffentlicht: (2024)
Exploring Traffic Crash Narratives in Jordan Using Text Mining Analytics
von: Jaradat, Shadi, et al.
Veröffentlicht: (2024)
von: Jaradat, Shadi, et al.
Veröffentlicht: (2024)
Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
von: Wei, Xinyu, et al.
Veröffentlicht: (2025)
von: Wei, Xinyu, et al.
Veröffentlicht: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
von: Du, Yifan, et al.
Veröffentlicht: (2023)
von: Du, Yifan, et al.
Veröffentlicht: (2023)
Exploring Visual Culture Awareness in GPT-4V: A Comprehensive Probing
von: Cao, Yong, et al.
Veröffentlicht: (2024)
von: Cao, Yong, et al.
Veröffentlicht: (2024)
TaxonRL: Reinforcement Learning with Intermediate Rewards for Interpretable Fine-Grained Visual Reasoning
von: von Klinski, Maximilian, et al.
Veröffentlicht: (2026)
von: von Klinski, Maximilian, et al.
Veröffentlicht: (2026)
FLUID: A Fine-Grained Lightweight Urban Signalized-Intersection Dataset of Dense Conflict Trajectories
von: Chen, Yiyang, et al.
Veröffentlicht: (2025)
von: Chen, Yiyang, et al.
Veröffentlicht: (2025)
Towards Equitable AI: Detecting Bias in Using Large Language Models for Marketing
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025)
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025)
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
von: Yu, Yongcan, et al.
Veröffentlicht: (2025)
von: Yu, Yongcan, et al.
Veröffentlicht: (2025)
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps
von: Feng, Sicheng, et al.
Veröffentlicht: (2025)
von: Feng, Sicheng, et al.
Veröffentlicht: (2025)
Expert Pyramid Tuning: Efficient Parameter Fine-Tuning for Expertise-Driven Task Allocation
von: Zhang, Jia-Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Jia-Chen, et al.
Veröffentlicht: (2026)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
von: Sun, Haoyuan, et al.
Veröffentlicht: (2025)
von: Sun, Haoyuan, et al.
Veröffentlicht: (2025)
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
von: Yanuka, Moran, et al.
Veröffentlicht: (2024)
von: Yanuka, Moran, et al.
Veröffentlicht: (2024)
Latent Visual Reasoning
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning
von: Song, Yingjin, et al.
Veröffentlicht: (2024)
von: Song, Yingjin, et al.
Veröffentlicht: (2024)
Learning Adaptive Reasoning Paths for Efficient Visual Reasoning
von: Huang, Yixu, et al.
Veröffentlicht: (2026)
von: Huang, Yixu, et al.
Veröffentlicht: (2026)
Evaluation of GPT-4o and GPT-4o-mini's Vision Capabilities for Compositional Analysis from Dried Solution Drops
von: Dangi, Deven B., et al.
Veröffentlicht: (2024)
von: Dangi, Deven B., et al.
Veröffentlicht: (2024)
Arbitration Failure, Not Perceptual Blindness: How Vision-Language Models Resolve Visual-Linguistic Conflicts
von: Nooralahzadeh, Farhad, et al.
Veröffentlicht: (2026)
von: Nooralahzadeh, Farhad, et al.
Veröffentlicht: (2026)
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
von: You, Haoxuan, et al.
Veröffentlicht: (2023)
von: You, Haoxuan, et al.
Veröffentlicht: (2023)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
von: Shen, Yifan, et al.
Veröffentlicht: (2025)
von: Shen, Yifan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Leveraging Large Language Models (LLMs) for Traffic Management at Urban Intersections: The Case of Mixed Traffic Scenarios
von: Masri, Sari, et al.
Veröffentlicht: (2024) -
Large Language Models (LLMs) as Traffic Control Systems at Urban Intersections: A New Paradigm
von: Masri, Sari, et al.
Veröffentlicht: (2024) -
HazardNet: A Small-Scale Vision Language Model for Real-Time Traffic Safety Detection at Edge Devices
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2025) -
Multimodal Large Language Models for Enhanced Traffic Safety: A Comprehensive Review and Future Trends
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2025) -
Using Multimodal Large Language Models for Automated Detection of Traffic Safety Critical Events
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2024)