Evaluation and Comparison of Visual Language Models for Transportation Engineering Problems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Prajapati, Sanjita, Singh, Tanu, Hegde, Chinmay, Chakraborty, Pranamesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Surveillance Video-Based Traffic Accident Detection Using Transformer Architecture
von: Singh, Tanu, et al.
Veröffentlicht: (2025)
von: Singh, Tanu, et al.
Veröffentlicht: (2025)
SGAP-Gaze: Scene Grid Attention Based Point-of-Gaze Estimation Network for Driver Gaze
von: Sharma, Pavan Kumar, et al.
Veröffentlicht: (2026)
von: Sharma, Pavan Kumar, et al.
Veröffentlicht: (2026)
FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis
von: Gavas, Ekta, et al.
Veröffentlicht: (2025)
von: Gavas, Ekta, et al.
Veröffentlicht: (2025)
A Review of Driver Gaze Estimation and Application in Gaze Behavior Understanding
von: Sharma, Pavan Kumar, et al.
Veröffentlicht: (2023)
von: Sharma, Pavan Kumar, et al.
Veröffentlicht: (2023)
DiffClean: Diffusion-based Makeup Removal for Accurate Age Estimation
von: Gavas, Ekta, et al.
Veröffentlicht: (2025)
von: Gavas, Ekta, et al.
Veröffentlicht: (2025)
FaceCloak: Learning to Protect Face Templates
von: Banerjee, Sudipta, et al.
Veröffentlicht: (2025)
von: Banerjee, Sudipta, et al.
Veröffentlicht: (2025)
Robust Concept Erasure Using Task Vectors
von: Pham, Minh, et al.
Veröffentlicht: (2024)
von: Pham, Minh, et al.
Veröffentlicht: (2024)
GOTCHA: Real-Time Video Deepfake Detection via Challenge-Response
von: Mittal, Govind, et al.
Veröffentlicht: (2022)
von: Mittal, Govind, et al.
Veröffentlicht: (2022)
Mitigating the Impact of Attribute Editing on Face Recognition
von: Banerjee, Sudipta, et al.
Veröffentlicht: (2024)
von: Banerjee, Sudipta, et al.
Veröffentlicht: (2024)
Vision-Language Models can Identify Distracted Driver Behavior from Naturalistic Videos
von: Hasan, Md Zahid, et al.
Veröffentlicht: (2023)
von: Hasan, Md Zahid, et al.
Veröffentlicht: (2023)
Leveraging Vision Language Models for Specialized Agricultural Tasks
von: Arshad, Muhammad Arbab, et al.
Veröffentlicht: (2024)
von: Arshad, Muhammad Arbab, et al.
Veröffentlicht: (2024)
RLS3: RL-Based Synthetic Sample Selection to Enhance Spatial Reasoning in Vision-Language Models for Indoor Autonomous Perception
von: Waite, Joshua R., et al.
Veröffentlicht: (2025)
von: Waite, Joshua R., et al.
Veröffentlicht: (2025)
Decorum: A Language-Based Approach For Style-Conditioned Synthesis of Indoor 3D Scenes
von: Marshall, Kelly O., et al.
Veröffentlicht: (2025)
von: Marshall, Kelly O., et al.
Veröffentlicht: (2025)
Visual Prompt Engineering for Vision Language Models in Radiology
von: Denner, Stefan, et al.
Veröffentlicht: (2024)
von: Denner, Stefan, et al.
Veröffentlicht: (2024)
ChartQA-X: Generating Explanations for Visual Chart Reasoning
von: Hegde, Shamanthak, et al.
Veröffentlicht: (2025)
von: Hegde, Shamanthak, et al.
Veröffentlicht: (2025)
When Visuals Aren't the Problem: Evaluating Vision-Language Models on Misleading Data Visualizations
von: Lalai, Harsh Nishant, et al.
Veröffentlicht: (2026)
von: Lalai, Harsh Nishant, et al.
Veröffentlicht: (2026)
When Are Concepts Erased From Diffusion Models?
von: Lu, Kevin, et al.
Veröffentlicht: (2025)
von: Lu, Kevin, et al.
Veröffentlicht: (2025)
Fast Certification of Vision-Language Models Using Incremental Randomized Smoothing
von: Nirala, A K, et al.
Veröffentlicht: (2023)
von: Nirala, A K, et al.
Veröffentlicht: (2023)
MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Space
von: Singh, Anshul, et al.
Veröffentlicht: (2025)
von: Singh, Anshul, et al.
Veröffentlicht: (2025)
SEAL: Semantic Aware Image Watermarking
von: Arabi, Kasra, et al.
Veröffentlicht: (2025)
von: Arabi, Kasra, et al.
Veröffentlicht: (2025)
SELECT: A Large-Scale Benchmark of Data Curation Strategies for Image Classification
von: Feuer, Benjamin, et al.
Veröffentlicht: (2024)
von: Feuer, Benjamin, et al.
Veröffentlicht: (2024)
Camera Perspective Transformation to Bird's Eye View via Spatial Transformer Model for Road Intersection Monitoring
von: Prajapati, Rukesh, et al.
Veröffentlicht: (2024)
von: Prajapati, Rukesh, et al.
Veröffentlicht: (2024)
Clean Evaluations on Contaminated Visual Language Models
von: Lu, Hongyuan, et al.
Veröffentlicht: (2024)
von: Lu, Hongyuan, et al.
Veröffentlicht: (2024)
Understanding and Evaluating Hallucinations in 3D Visual Language Models
von: Peng, Ruiying, et al.
Veröffentlicht: (2025)
von: Peng, Ruiying, et al.
Veröffentlicht: (2025)
Attention-Gated Convolutional Networks for Scanner-Agnostic Quality Assessment
von: Bakhale, Chinmay, et al.
Veröffentlicht: (2026)
von: Bakhale, Chinmay, et al.
Veröffentlicht: (2026)
Slice-100K: A Multimodal Dataset for Extrusion-based 3D Printing
von: Jignasu, Anushrut, et al.
Veröffentlicht: (2024)
von: Jignasu, Anushrut, et al.
Veröffentlicht: (2024)
Comparison Visual Instruction Tuning
von: Lin, Wei, et al.
Veröffentlicht: (2024)
von: Lin, Wei, et al.
Veröffentlicht: (2024)
The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?
von: Ghosh, Samrajnee, et al.
Veröffentlicht: (2025)
von: Ghosh, Samrajnee, et al.
Veröffentlicht: (2025)
Hidden in the Noise: Two-Stage Robust Watermarking for Images
von: Arabi, Kasra, et al.
Veröffentlicht: (2024)
von: Arabi, Kasra, et al.
Veröffentlicht: (2024)
How Reasoning Influences Intersectional Biases in Vision Language Models
von: Desai, Adit, et al.
Veröffentlicht: (2025)
von: Desai, Adit, et al.
Veröffentlicht: (2025)
TerraIncognita: A Dynamic Benchmark for Species Discovery Using Frontier Models
von: Chiranjeevi, Shivani, et al.
Veröffentlicht: (2025)
von: Chiranjeevi, Shivani, et al.
Veröffentlicht: (2025)
CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models
von: Cai, Jie, et al.
Veröffentlicht: (2025)
von: Cai, Jie, et al.
Veröffentlicht: (2025)
Language Model as Visual Explainer
von: Yang, Xingyi, et al.
Veröffentlicht: (2024)
von: Yang, Xingyi, et al.
Veröffentlicht: (2024)
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
von: Leng, Sicong, et al.
Veröffentlicht: (2024)
von: Leng, Sicong, et al.
Veröffentlicht: (2024)
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
von: Xu, Weiye, et al.
Veröffentlicht: (2025)
von: Xu, Weiye, et al.
Veröffentlicht: (2025)
Autonomous Crack Detection using Deep Learning on Synthetic Thermogram Datasets
von: Pimpalkhare, Chinmay Makarand, et al.
Veröffentlicht: (2024)
von: Pimpalkhare, Chinmay Makarand, et al.
Veröffentlicht: (2024)
Optimal Transport Aggregation for Visual Place Recognition
von: Izquierdo, Sergio, et al.
Veröffentlicht: (2023)
von: Izquierdo, Sergio, et al.
Veröffentlicht: (2023)
The Visual Iconicity Challenge: Evaluating Vision-Language Models on Sign Language Form-Meaning Mapping
von: Keleş, Onur, et al.
Veröffentlicht: (2025)
von: Keleş, Onur, et al.
Veröffentlicht: (2025)
Towards Open-ended Visual Quality Comparison
von: Wu, Haoning, et al.
Veröffentlicht: (2024)
von: Wu, Haoning, et al.
Veröffentlicht: (2024)
REVEAL -- Reasoning and Evaluation of Visual Evidence through Aligned Language
von: Praharaj, Ipsita, et al.
Veröffentlicht: (2025)
von: Praharaj, Ipsita, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Surveillance Video-Based Traffic Accident Detection Using Transformer Architecture
von: Singh, Tanu, et al.
Veröffentlicht: (2025) -
SGAP-Gaze: Scene Grid Attention Based Point-of-Gaze Estimation Network for Driver Gaze
von: Sharma, Pavan Kumar, et al.
Veröffentlicht: (2026) -
FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis
von: Gavas, Ekta, et al.
Veröffentlicht: (2025) -
A Review of Driver Gaze Estimation and Application in Gaze Behavior Understanding
von: Sharma, Pavan Kumar, et al.
Veröffentlicht: (2023) -
DiffClean: Diffusion-based Makeup Removal for Accurate Age Estimation
von: Gavas, Ekta, et al.
Veröffentlicht: (2025)