Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Elhenawy, Mohammed, Jaradat, Shadi, Alhadidi, Taqwa I., Ashqar, Huthaifa I., Jaber, Ahmed, Rakotonirainy, Andry, Tami, Mohammad Abu |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding
par: Elhenawy, Mohammed, et autres
Publié: (2025)
par: Elhenawy, Mohammed, et autres
Publié: (2025)
Object Detection using Oriented Window Learning Vi-sion Transformer: Roadway Assets Recognition
par: Alhadidi, Taqwa, et autres
Publié: (2024)
par: Alhadidi, Taqwa, et autres
Publié: (2024)
Visual Reasoning and Multi-Agent Approach in Multimodal Large Language Models (MLLMs): Solving TSP and mTSP Combinatorial Challenges
par: Elhenawy, Mohammed, et autres
Publié: (2024)
par: Elhenawy, Mohammed, et autres
Publié: (2024)
Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing
par: Ashqar, Huthaifa I., et autres
Publié: (2024)
par: Ashqar, Huthaifa I., et autres
Publié: (2024)
Eyeballing Combinatorial Problems: A Case Study of Using Multimodal Large Language Models to Solve Traveling Salesman Problems
par: Elhenawy, Mohammed, et autres
Publié: (2024)
par: Elhenawy, Mohammed, et autres
Publié: (2024)
Enhancing Pavement Crack Classification with Bidirectional Cascaded Neural Networks
par: Alhadidi, Taqwa I., et autres
Publié: (2025)
par: Alhadidi, Taqwa I., et autres
Publié: (2025)
Exploring Traffic Crash Narratives in Jordan Using Text Mining Analytics
par: Jaradat, Shadi, et autres
Publié: (2024)
par: Jaradat, Shadi, et autres
Publié: (2024)
Using Multimodal Large Language Models for Automated Detection of Traffic Safety Critical Events
par: Tami, Mohammad Abu, et autres
Publié: (2024)
par: Tami, Mohammad Abu, et autres
Publié: (2024)
The Use of Multimodal Large Language Models to Detect Objects from Thermal Images: Transportation Applications
par: Ashqar, Huthaifa I., et autres
Publié: (2024)
par: Ashqar, Huthaifa I., et autres
Publié: (2024)
Multimodal Large Language Models for Enhanced Traffic Safety: A Comprehensive Review and Future Trends
par: Tami, Mohammad Abu, et autres
Publié: (2025)
par: Tami, Mohammad Abu, et autres
Publié: (2025)
Automated Question Generation for Science Tests in Arabic Language Using NLP Techniques
par: Tami, Mohammad, et autres
Publié: (2024)
par: Tami, Mohammad, et autres
Publié: (2024)
HazardNet: A Small-Scale Vision Language Model for Real-Time Traffic Safety Detection at Edge Devices
par: Tami, Mohammad Abu, et autres
Publié: (2025)
par: Tami, Mohammad Abu, et autres
Publié: (2025)
Large Language Models (LLMs) as Traffic Control Systems at Urban Intersections: A New Paradigm
par: Masri, Sari, et autres
Publié: (2024)
par: Masri, Sari, et autres
Publié: (2024)
Leveraging Large Language Models (LLMs) for Traffic Management at Urban Intersections: The Case of Mixed Traffic Scenarios
par: Masri, Sari, et autres
Publié: (2024)
par: Masri, Sari, et autres
Publié: (2024)
Automated Pavement Cracks Detection and Classification Using Deep Learning
par: Nafaa, Selvia, et autres
Publié: (2024)
par: Nafaa, Selvia, et autres
Publié: (2024)
Visual Reasoning at Urban Intersections: FineTuning GPT-4o for Traffic Conflict Detection
par: Masri, Sari, et autres
Publié: (2025)
par: Masri, Sari, et autres
Publié: (2025)
Question-Answering (QA) Model for a Personalized Learning Assistant for Arabic Language
par: Sammoudi, Mohammad, et autres
Publié: (2024)
par: Sammoudi, Mohammad, et autres
Publié: (2024)
Advancing Roadway Sign Detection with YOLO Models and Transfer Learning
par: Nafaa, Selvia, et autres
Publié: (2024)
par: Nafaa, Selvia, et autres
Publié: (2024)
Transformer Models in Education: Summarizing Science Textbooks with AraBART, MT5, AraT5, and mBART
par: Masri, Sari, et autres
Publié: (2024)
par: Masri, Sari, et autres
Publié: (2024)
Towards Equitable AI: Detecting Bias in Using Large Language Models for Marketing
par: Yilmaz, Berk, et autres
Publié: (2025)
par: Yilmaz, Berk, et autres
Publié: (2025)
Impact of Connected and Automated Vehicles on Transport Injustices
par: Martinez-Buelvas, Laura, et autres
Publié: (2024)
par: Martinez-Buelvas, Laura, et autres
Publié: (2024)
Enhancing Mathematics Learning for Hard-of-Hearing Students Through Real-Time Palestinian Sign Language Recognition: A New Dataset
par: Khandaqji, Fidaa, et autres
Publié: (2025)
par: Khandaqji, Fidaa, et autres
Publié: (2025)
X-Blocks: Linguistic Building Blocks of Natural Language Explanations for Automated Vehicles
par: Zadeh, Ashkan Y., et autres
Publié: (2026)
par: Zadeh, Ashkan Y., et autres
Publié: (2026)
Leveraging Explainable AI for LLM Text Attribution: Differentiating Human-Written and Multiple LLMs-Generated Text
par: Najjar, Ayat, et autres
Publié: (2025)
par: Najjar, Ayat, et autres
Publié: (2025)
Malware Classification from Memory Dumps Using Machine Learning, Transformers, and Large Language Models
par: Dweib, Areej, et autres
Publié: (2025)
par: Dweib, Areej, et autres
Publié: (2025)
Evaluation and Optimization of Adaptive Cruise Control in Autonomous Vehicles using the CARLA Simulator: A Study on Performance under Wet and Dry Weather Conditions
par: Al-Hindaw, Roza, et autres
Publié: (2024)
par: Al-Hindaw, Roza, et autres
Publié: (2024)
Network Traffic Classification Using Machine Learning, Transformer, and Large Language Models
par: Antari, Ahmad, et autres
Publié: (2025)
par: Antari, Ahmad, et autres
Publié: (2025)
Detecting AI-Generated Text in Educational Content: Leveraging Machine Learning and Explainable AI for Academic Integrity
par: Najjar, Ayat A., et autres
Publié: (2025)
par: Najjar, Ayat A., et autres
Publié: (2025)
When to Commute During the COVID-19 Pandemic and Beyond: Analysis of Traffic Crashes in Washington, D.C
par: Choi, Joanne, et autres
Publié: (2024)
par: Choi, Joanne, et autres
Publié: (2024)
Flood Risk Assessment of the National Harbor at Maryland, United States
par: Negussie, Neftalem, et autres
Publié: (2024)
par: Negussie, Neftalem, et autres
Publié: (2024)
Influential Factors in Increasing an Amazon products Sales Rank
par: Chen, Ben, et autres
Publié: (2024)
par: Chen, Ben, et autres
Publié: (2024)
Exploring Combinatorial Problem Solving with Large Language Models: A Case Study on the Travelling Salesman Problem Using GPT-3.5 Turbo
par: Masoud, Mahmoud, et autres
Publié: (2024)
par: Masoud, Mahmoud, et autres
Publié: (2024)
Ride-sharing Determinants: Spatial and Spatio-temporal Bayesian Analysis for Chicago Service in 2022
par: Elkhouly, Mohamed, et autres
Publié: (2024)
par: Elkhouly, Mohamed, et autres
Publié: (2024)
The Effect of Funding on Student Achievement: Evidence from District of Columbia, Virginia, and Maryland
par: Raabe, Adam, et autres
Publié: (2024)
par: Raabe, Adam, et autres
Publié: (2024)
Analysis of Droughts and Their Intensities in California from 2000 to 2020
par: Ujjwal, et autres
Publié: (2024)
par: Ujjwal, et autres
Publié: (2024)
An Eye Gaze Heatmap Analysis of Uncertainty Head-Up Display Designs for Conditional Automated Driving
par: Gerber, Michael A., et autres
Publié: (2024)
par: Gerber, Michael A., et autres
Publié: (2024)
The Impact of Medicaid Expansion on Medicare Quality Measures
par: Algrain, Hala, et autres
Publié: (2024)
par: Algrain, Hala, et autres
Publié: (2024)
Small Talk, Big Impact? LLM-based Conversational Agents to Mitigate Passive Fatigue in Conditional Automated Driving
par: Cockram, Lewis, et autres
Publié: (2025)
par: Cockram, Lewis, et autres
Publié: (2025)
Identifying Economic Factors Affecting Unemployment Rates in the United States
par: Green, Alrick, et autres
Publié: (2024)
par: Green, Alrick, et autres
Publié: (2024)
UniFlow: Zero-Shot LiDAR Scene Flow for Autonomous Vehicles
par: Li, Siyi, et autres
Publié: (2025)
par: Li, Siyi, et autres
Publié: (2025)
Documents similaires
-
Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding
par: Elhenawy, Mohammed, et autres
Publié: (2025) -
Object Detection using Oriented Window Learning Vi-sion Transformer: Roadway Assets Recognition
par: Alhadidi, Taqwa, et autres
Publié: (2024) -
Visual Reasoning and Multi-Agent Approach in Multimodal Large Language Models (MLLMs): Solving TSP and mTSP Combinatorial Challenges
par: Elhenawy, Mohammed, et autres
Publié: (2024) -
Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing
par: Ashqar, Huthaifa I., et autres
Publié: (2024) -
Eyeballing Combinatorial Problems: A Case Study of Using Multimodal Large Language Models to Solve Traveling Salesman Problems
par: Elhenawy, Mohammed, et autres
Publié: (2024)