The Use of Multimodal Large Language Models to Detect Objects from Thermal Images: Transportation Applications
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ashqar, Huthaifa I., Alhadidi, Taqwa I., Elhenawy, Mohammed, Khanfar, Nour O. |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing
par: Ashqar, Huthaifa I., et autres
Publié: (2024)
par: Ashqar, Huthaifa I., et autres
Publié: (2024)
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
par: Elhenawy, Mohammed, et autres
Publié: (2025)
par: Elhenawy, Mohammed, et autres
Publié: (2025)
Using Multimodal Large Language Models for Automated Detection of Traffic Safety Critical Events
par: Tami, Mohammad Abu, et autres
Publié: (2024)
par: Tami, Mohammad Abu, et autres
Publié: (2024)
Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding
par: Elhenawy, Mohammed, et autres
Publié: (2025)
par: Elhenawy, Mohammed, et autres
Publié: (2025)
Object Detection using Oriented Window Learning Vi-sion Transformer: Roadway Assets Recognition
par: Alhadidi, Taqwa, et autres
Publié: (2024)
par: Alhadidi, Taqwa, et autres
Publié: (2024)
Multimodal Large Language Models for Enhanced Traffic Safety: A Comprehensive Review and Future Trends
par: Tami, Mohammad Abu, et autres
Publié: (2025)
par: Tami, Mohammad Abu, et autres
Publié: (2025)
Visual Reasoning at Urban Intersections: FineTuning GPT-4o for Traffic Conflict Detection
par: Masri, Sari, et autres
Publié: (2025)
par: Masri, Sari, et autres
Publié: (2025)
Leveraging Large Language Models (LLMs) for Traffic Management at Urban Intersections: The Case of Mixed Traffic Scenarios
par: Masri, Sari, et autres
Publié: (2024)
par: Masri, Sari, et autres
Publié: (2024)
Automated Pavement Cracks Detection and Classification Using Deep Learning
par: Nafaa, Selvia, et autres
Publié: (2024)
par: Nafaa, Selvia, et autres
Publié: (2024)
Advancing Roadway Sign Detection with YOLO Models and Transfer Learning
par: Nafaa, Selvia, et autres
Publié: (2024)
par: Nafaa, Selvia, et autres
Publié: (2024)
HazardNet: A Small-Scale Vision Language Model for Real-Time Traffic Safety Detection at Edge Devices
par: Tami, Mohammad Abu, et autres
Publié: (2025)
par: Tami, Mohammad Abu, et autres
Publié: (2025)
Automated Question Generation for Science Tests in Arabic Language Using NLP Techniques
par: Tami, Mohammad, et autres
Publié: (2024)
par: Tami, Mohammad, et autres
Publié: (2024)
Enhancing Pavement Crack Classification with Bidirectional Cascaded Neural Networks
par: Alhadidi, Taqwa I., et autres
Publié: (2025)
par: Alhadidi, Taqwa I., et autres
Publié: (2025)
Large Language Models (LLMs) as Traffic Control Systems at Urban Intersections: A New Paradigm
par: Masri, Sari, et autres
Publié: (2024)
par: Masri, Sari, et autres
Publié: (2024)
Exploring Traffic Crash Narratives in Jordan Using Text Mining Analytics
par: Jaradat, Shadi, et autres
Publié: (2024)
par: Jaradat, Shadi, et autres
Publié: (2024)
Question-Answering (QA) Model for a Personalized Learning Assistant for Arabic Language
par: Sammoudi, Mohammad, et autres
Publié: (2024)
par: Sammoudi, Mohammad, et autres
Publié: (2024)
Towards Equitable AI: Detecting Bias in Using Large Language Models for Marketing
par: Yilmaz, Berk, et autres
Publié: (2025)
par: Yilmaz, Berk, et autres
Publié: (2025)
Enhancing Mathematics Learning for Hard-of-Hearing Students Through Real-Time Palestinian Sign Language Recognition: A New Dataset
par: Khandaqji, Fidaa, et autres
Publié: (2025)
par: Khandaqji, Fidaa, et autres
Publié: (2025)
Leveraging Explainable AI for LLM Text Attribution: Differentiating Human-Written and Multiple LLMs-Generated Text
par: Najjar, Ayat, et autres
Publié: (2025)
par: Najjar, Ayat, et autres
Publié: (2025)
Visual Reasoning and Multi-Agent Approach in Multimodal Large Language Models (MLLMs): Solving TSP and mTSP Combinatorial Challenges
par: Elhenawy, Mohammed, et autres
Publié: (2024)
par: Elhenawy, Mohammed, et autres
Publié: (2024)
Detecting AI-Generated Text in Educational Content: Leveraging Machine Learning and Explainable AI for Academic Integrity
par: Najjar, Ayat A., et autres
Publié: (2025)
par: Najjar, Ayat A., et autres
Publié: (2025)
Transformer Models in Education: Summarizing Science Textbooks with AraBART, MT5, AraT5, and mBART
par: Masri, Sari, et autres
Publié: (2024)
par: Masri, Sari, et autres
Publié: (2024)
MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms
par: Jin, Yiqiao, et autres
Publié: (2024)
par: Jin, Yiqiao, et autres
Publié: (2024)
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
par: Qi, Peng, et autres
Publié: (2024)
par: Qi, Peng, et autres
Publié: (2024)
Large Language Models and Provenance Metadata for Determining the Relevance of Images and Videos in News Stories
par: Peterka, Tomas, et autres
Publié: (2025)
par: Peterka, Tomas, et autres
Publié: (2025)
Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images
par: Fraser, Kathleen C., et autres
Publié: (2024)
par: Fraser, Kathleen C., et autres
Publié: (2024)
ERIT Lightweight Multimodal Dataset for Elderly Emotion Recognition and Multimodal Fusion Evaluation
par: Frieske, Rita, et autres
Publié: (2024)
par: Frieske, Rita, et autres
Publié: (2024)
Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
par: Yan, Zehong, et autres
Publié: (2025)
par: Yan, Zehong, et autres
Publié: (2025)
Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
par: Sapkota, Ranjan, et autres
Publié: (2025)
par: Sapkota, Ranjan, et autres
Publié: (2025)
Restoring Ancient Ideograph: A Multimodal Multitask Neural Network Approach
par: Duan, Siyu, et autres
Publié: (2024)
par: Duan, Siyu, et autres
Publié: (2024)
Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
par: Si, Chenglei, et autres
Publié: (2024)
par: Si, Chenglei, et autres
Publié: (2024)
Describing Differences in Image Sets with Natural Language
par: Dunlap, Lisa, et autres
Publié: (2023)
par: Dunlap, Lisa, et autres
Publié: (2023)
Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags
par: Qi, Daiqing, et autres
Publié: (2024)
par: Qi, Daiqing, et autres
Publié: (2024)
Semantic and Expressive Variation in Image Captions Across Languages
par: Ye, Andre, et autres
Publié: (2023)
par: Ye, Andre, et autres
Publié: (2023)
Two Stage Context Learning with Large Language Models for Multimodal Stance Detection on Climate Change
par: Pangtey, Lata, et autres
Publié: (2025)
par: Pangtey, Lata, et autres
Publié: (2025)
D3G: Diverse Demographic Data Generation Increases Zero-Shot Image Classification Accuracy within Multimodal Models
par: Hickmon, Javon
Publié: (2025)
par: Hickmon, Javon
Publié: (2025)
Stable Signer: Hierarchical Sign Language Generative Model
par: Fang, Sen, et autres
Publié: (2025)
par: Fang, Sen, et autres
Publié: (2025)
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation
par: Jha, Akshita, et autres
Publié: (2024)
par: Jha, Akshita, et autres
Publié: (2024)
The Impact of Image Resolution on Biomedical Multimodal Large Language Models
par: Chen, Liangyu, et autres
Publié: (2025)
par: Chen, Liangyu, et autres
Publié: (2025)
Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
par: Nazi, Zabir Al, et autres
Publié: (2025)
par: Nazi, Zabir Al, et autres
Publié: (2025)
Documents similaires
-
Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing
par: Ashqar, Huthaifa I., et autres
Publié: (2024) -
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
par: Elhenawy, Mohammed, et autres
Publié: (2025) -
Using Multimodal Large Language Models for Automated Detection of Traffic Safety Critical Events
par: Tami, Mohammad Abu, et autres
Publié: (2024) -
Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding
par: Elhenawy, Mohammed, et autres
Publié: (2025) -
Object Detection using Oriented Window Learning Vi-sion Transformer: Roadway Assets Recognition
par: Alhadidi, Taqwa, et autres
Publié: (2024)