Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ashqar, Huthaifa I., Jaber, Ahmed, Alhadidi, Taqwa I., Elhenawy, Mohammed |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Use of Multimodal Large Language Models to Detect Objects from Thermal Images: Transportation Applications
von: Ashqar, Huthaifa I., et al.
Veröffentlicht: (2024)
von: Ashqar, Huthaifa I., et al.
Veröffentlicht: (2024)
Object Detection using Oriented Window Learning Vi-sion Transformer: Roadway Assets Recognition
von: Alhadidi, Taqwa, et al.
Veröffentlicht: (2024)
von: Alhadidi, Taqwa, et al.
Veröffentlicht: (2024)
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
Multimodal Large Language Models for Enhanced Traffic Safety: A Comprehensive Review and Future Trends
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2025)
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2025)
Enhancing Pavement Crack Classification with Bidirectional Cascaded Neural Networks
von: Alhadidi, Taqwa I., et al.
Veröffentlicht: (2025)
von: Alhadidi, Taqwa I., et al.
Veröffentlicht: (2025)
Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
Visual Reasoning at Urban Intersections: FineTuning GPT-4o for Traffic Conflict Detection
von: Masri, Sari, et al.
Veröffentlicht: (2025)
von: Masri, Sari, et al.
Veröffentlicht: (2025)
HazardNet: A Small-Scale Vision Language Model for Real-Time Traffic Safety Detection at Edge Devices
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2025)
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2025)
Using Multimodal Large Language Models for Automated Detection of Traffic Safety Critical Events
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2024)
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2024)
Exploring Traffic Crash Narratives in Jordan Using Text Mining Analytics
von: Jaradat, Shadi, et al.
Veröffentlicht: (2024)
von: Jaradat, Shadi, et al.
Veröffentlicht: (2024)
Advancing Roadway Sign Detection with YOLO Models and Transfer Learning
von: Nafaa, Selvia, et al.
Veröffentlicht: (2024)
von: Nafaa, Selvia, et al.
Veröffentlicht: (2024)
Visual Reasoning and Multi-Agent Approach in Multimodal Large Language Models (MLLMs): Solving TSP and mTSP Combinatorial Challenges
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2024)
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2024)
Automated Question Generation for Science Tests in Arabic Language Using NLP Techniques
von: Tami, Mohammad, et al.
Veröffentlicht: (2024)
von: Tami, Mohammad, et al.
Veröffentlicht: (2024)
Automated Pavement Cracks Detection and Classification Using Deep Learning
von: Nafaa, Selvia, et al.
Veröffentlicht: (2024)
von: Nafaa, Selvia, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models (LLMs) for Traffic Management at Urban Intersections: The Case of Mixed Traffic Scenarios
von: Masri, Sari, et al.
Veröffentlicht: (2024)
von: Masri, Sari, et al.
Veröffentlicht: (2024)
Large Language Models (LLMs) as Traffic Control Systems at Urban Intersections: A New Paradigm
von: Masri, Sari, et al.
Veröffentlicht: (2024)
von: Masri, Sari, et al.
Veröffentlicht: (2024)
Question-Answering (QA) Model for a Personalized Learning Assistant for Arabic Language
von: Sammoudi, Mohammad, et al.
Veröffentlicht: (2024)
von: Sammoudi, Mohammad, et al.
Veröffentlicht: (2024)
Eyeballing Combinatorial Problems: A Case Study of Using Multimodal Large Language Models to Solve Traveling Salesman Problems
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2024)
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2024)
Towards Equitable AI: Detecting Bias in Using Large Language Models for Marketing
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025)
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025)
Transformer Models in Education: Summarizing Science Textbooks with AraBART, MT5, AraT5, and mBART
von: Masri, Sari, et al.
Veröffentlicht: (2024)
von: Masri, Sari, et al.
Veröffentlicht: (2024)
Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
Enhancing Mathematics Learning for Hard-of-Hearing Students Through Real-Time Palestinian Sign Language Recognition: A New Dataset
von: Khandaqji, Fidaa, et al.
Veröffentlicht: (2025)
von: Khandaqji, Fidaa, et al.
Veröffentlicht: (2025)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge Alignment
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
CODIS: Benchmarking Context-Dependent Visual Comprehension for Multimodal Large Language Models
von: Luo, Fuwen, et al.
Veröffentlicht: (2024)
von: Luo, Fuwen, et al.
Veröffentlicht: (2024)
Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
von: Huang, Yiran, et al.
Veröffentlicht: (2025)
von: Huang, Yiran, et al.
Veröffentlicht: (2025)
Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags
von: Qi, Daiqing, et al.
Veröffentlicht: (2024)
von: Qi, Daiqing, et al.
Veröffentlicht: (2024)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability
von: Wu, Daiqing, et al.
Veröffentlicht: (2025)
von: Wu, Daiqing, et al.
Veröffentlicht: (2025)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2025)
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2025)
Mask What Matters: Mitigating Object Hallucinations in Multimodal Large Language Models with Object-Aligned Visual Contrastive Decoding
von: Chen, Boqi, et al.
Veröffentlicht: (2026)
von: Chen, Boqi, et al.
Veröffentlicht: (2026)
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025)
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025)
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
VidCoM: Fast Video Comprehension through Large Language Models with Multimodal Tools
von: Qi, Ji, et al.
Veröffentlicht: (2023)
von: Qi, Ji, et al.
Veröffentlicht: (2023)
UniChange: Unifying Change Detection with Multimodal Large Language Model
von: Zhang, Xu, et al.
Veröffentlicht: (2025)
von: Zhang, Xu, et al.
Veröffentlicht: (2025)
Network Traffic Classification Using Machine Learning, Transformer, and Large Language Models
von: Antari, Ahmad, et al.
Veröffentlicht: (2025)
von: Antari, Ahmad, et al.
Veröffentlicht: (2025)
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Use of Multimodal Large Language Models to Detect Objects from Thermal Images: Transportation Applications
von: Ashqar, Huthaifa I., et al.
Veröffentlicht: (2024) -
Object Detection using Oriented Window Learning Vi-sion Transformer: Roadway Assets Recognition
von: Alhadidi, Taqwa, et al.
Veröffentlicht: (2024) -
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025) -
Multimodal Large Language Models for Enhanced Traffic Safety: A Comprehensive Review and Future Trends
von: Tami, Mohammad Abu, et al.
Veröffentlicht: (2025) -
Enhancing Pavement Crack Classification with Bidirectional Cascaded Neural Networks
von: Alhadidi, Taqwa I., et al.
Veröffentlicht: (2025)