The Use of Multimodal Large Language Models to Detect Objects from Thermal Images: Transportation Applications
Fuente:
arXiv
Saved in:
| Main Authors: | Ashqar, Huthaifa I., Alhadidi, Taqwa I., Elhenawy, Mohammed, Khanfar, Nour O. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing
by: Ashqar, Huthaifa I., et al.
Published: (2024)
by: Ashqar, Huthaifa I., et al.
Published: (2024)
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
by: Elhenawy, Mohammed, et al.
Published: (2025)
by: Elhenawy, Mohammed, et al.
Published: (2025)
Using Multimodal Large Language Models for Automated Detection of Traffic Safety Critical Events
by: Tami, Mohammad Abu, et al.
Published: (2024)
by: Tami, Mohammad Abu, et al.
Published: (2024)
Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding
by: Elhenawy, Mohammed, et al.
Published: (2025)
by: Elhenawy, Mohammed, et al.
Published: (2025)
Object Detection using Oriented Window Learning Vi-sion Transformer: Roadway Assets Recognition
by: Alhadidi, Taqwa, et al.
Published: (2024)
by: Alhadidi, Taqwa, et al.
Published: (2024)
Multimodal Large Language Models for Enhanced Traffic Safety: A Comprehensive Review and Future Trends
by: Tami, Mohammad Abu, et al.
Published: (2025)
by: Tami, Mohammad Abu, et al.
Published: (2025)
Visual Reasoning at Urban Intersections: FineTuning GPT-4o for Traffic Conflict Detection
by: Masri, Sari, et al.
Published: (2025)
by: Masri, Sari, et al.
Published: (2025)
Leveraging Large Language Models (LLMs) for Traffic Management at Urban Intersections: The Case of Mixed Traffic Scenarios
by: Masri, Sari, et al.
Published: (2024)
by: Masri, Sari, et al.
Published: (2024)
Automated Pavement Cracks Detection and Classification Using Deep Learning
by: Nafaa, Selvia, et al.
Published: (2024)
by: Nafaa, Selvia, et al.
Published: (2024)
Advancing Roadway Sign Detection with YOLO Models and Transfer Learning
by: Nafaa, Selvia, et al.
Published: (2024)
by: Nafaa, Selvia, et al.
Published: (2024)
HazardNet: A Small-Scale Vision Language Model for Real-Time Traffic Safety Detection at Edge Devices
by: Tami, Mohammad Abu, et al.
Published: (2025)
by: Tami, Mohammad Abu, et al.
Published: (2025)
Automated Question Generation for Science Tests in Arabic Language Using NLP Techniques
by: Tami, Mohammad, et al.
Published: (2024)
by: Tami, Mohammad, et al.
Published: (2024)
Enhancing Pavement Crack Classification with Bidirectional Cascaded Neural Networks
by: Alhadidi, Taqwa I., et al.
Published: (2025)
by: Alhadidi, Taqwa I., et al.
Published: (2025)
Large Language Models (LLMs) as Traffic Control Systems at Urban Intersections: A New Paradigm
by: Masri, Sari, et al.
Published: (2024)
by: Masri, Sari, et al.
Published: (2024)
Exploring Traffic Crash Narratives in Jordan Using Text Mining Analytics
by: Jaradat, Shadi, et al.
Published: (2024)
by: Jaradat, Shadi, et al.
Published: (2024)
Question-Answering (QA) Model for a Personalized Learning Assistant for Arabic Language
by: Sammoudi, Mohammad, et al.
Published: (2024)
by: Sammoudi, Mohammad, et al.
Published: (2024)
Towards Equitable AI: Detecting Bias in Using Large Language Models for Marketing
by: Yilmaz, Berk, et al.
Published: (2025)
by: Yilmaz, Berk, et al.
Published: (2025)
Enhancing Mathematics Learning for Hard-of-Hearing Students Through Real-Time Palestinian Sign Language Recognition: A New Dataset
by: Khandaqji, Fidaa, et al.
Published: (2025)
by: Khandaqji, Fidaa, et al.
Published: (2025)
Leveraging Explainable AI for LLM Text Attribution: Differentiating Human-Written and Multiple LLMs-Generated Text
by: Najjar, Ayat, et al.
Published: (2025)
by: Najjar, Ayat, et al.
Published: (2025)
Visual Reasoning and Multi-Agent Approach in Multimodal Large Language Models (MLLMs): Solving TSP and mTSP Combinatorial Challenges
by: Elhenawy, Mohammed, et al.
Published: (2024)
by: Elhenawy, Mohammed, et al.
Published: (2024)
Detecting AI-Generated Text in Educational Content: Leveraging Machine Learning and Explainable AI for Academic Integrity
by: Najjar, Ayat A., et al.
Published: (2025)
by: Najjar, Ayat A., et al.
Published: (2025)
Transformer Models in Education: Summarizing Science Textbooks with AraBART, MT5, AraT5, and mBART
by: Masri, Sari, et al.
Published: (2024)
by: Masri, Sari, et al.
Published: (2024)
MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms
by: Jin, Yiqiao, et al.
Published: (2024)
by: Jin, Yiqiao, et al.
Published: (2024)
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
by: Qi, Peng, et al.
Published: (2024)
by: Qi, Peng, et al.
Published: (2024)
Large Language Models and Provenance Metadata for Determining the Relevance of Images and Videos in News Stories
by: Peterka, Tomas, et al.
Published: (2025)
by: Peterka, Tomas, et al.
Published: (2025)
Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images
by: Fraser, Kathleen C., et al.
Published: (2024)
by: Fraser, Kathleen C., et al.
Published: (2024)
ERIT Lightweight Multimodal Dataset for Elderly Emotion Recognition and Multimodal Fusion Evaluation
by: Frieske, Rita, et al.
Published: (2024)
by: Frieske, Rita, et al.
Published: (2024)
Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
by: Yan, Zehong, et al.
Published: (2025)
by: Yan, Zehong, et al.
Published: (2025)
Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Restoring Ancient Ideograph: A Multimodal Multitask Neural Network Approach
by: Duan, Siyu, et al.
Published: (2024)
by: Duan, Siyu, et al.
Published: (2024)
Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
by: Si, Chenglei, et al.
Published: (2024)
by: Si, Chenglei, et al.
Published: (2024)
Describing Differences in Image Sets with Natural Language
by: Dunlap, Lisa, et al.
Published: (2023)
by: Dunlap, Lisa, et al.
Published: (2023)
Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags
by: Qi, Daiqing, et al.
Published: (2024)
by: Qi, Daiqing, et al.
Published: (2024)
Semantic and Expressive Variation in Image Captions Across Languages
by: Ye, Andre, et al.
Published: (2023)
by: Ye, Andre, et al.
Published: (2023)
Two Stage Context Learning with Large Language Models for Multimodal Stance Detection on Climate Change
by: Pangtey, Lata, et al.
Published: (2025)
by: Pangtey, Lata, et al.
Published: (2025)
D3G: Diverse Demographic Data Generation Increases Zero-Shot Image Classification Accuracy within Multimodal Models
by: Hickmon, Javon
Published: (2025)
by: Hickmon, Javon
Published: (2025)
Stable Signer: Hierarchical Sign Language Generative Model
by: Fang, Sen, et al.
Published: (2025)
by: Fang, Sen, et al.
Published: (2025)
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation
by: Jha, Akshita, et al.
Published: (2024)
by: Jha, Akshita, et al.
Published: (2024)
The Impact of Image Resolution on Biomedical Multimodal Large Language Models
by: Chen, Liangyu, et al.
Published: (2025)
by: Chen, Liangyu, et al.
Published: (2025)
Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
by: Nazi, Zabir Al, et al.
Published: (2025)
by: Nazi, Zabir Al, et al.
Published: (2025)
Similar Items
-
Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing
by: Ashqar, Huthaifa I., et al.
Published: (2024) -
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
by: Elhenawy, Mohammed, et al.
Published: (2025) -
Using Multimodal Large Language Models for Automated Detection of Traffic Safety Critical Events
by: Tami, Mohammad Abu, et al.
Published: (2024) -
Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding
by: Elhenawy, Mohammed, et al.
Published: (2025) -
Object Detection using Oriented Window Learning Vi-sion Transformer: Roadway Assets Recognition
by: Alhadidi, Taqwa, et al.
Published: (2024)