VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Younggun, Abdelrahman, Ahmed S., Abdel-Aty, Mohamed |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VRU-CIPI: Crossing Intention Prediction at Intersections for Improving Vulnerable Road Users Safety
di: Abdelrahman, Ahmed S., et al.
Pubblicazione: (2025)
di: Abdelrahman, Ahmed S., et al.
Pubblicazione: (2025)
Video-to-Text Pedestrian Monitoring (VTPM): Leveraging Computer Vision and Large Language Models for Privacy-Preserve Pedestrian Activity Monitoring at Intersections
di: Abdelrahman, Ahmed S., et al.
Pubblicazione: (2024)
di: Abdelrahman, Ahmed S., et al.
Pubblicazione: (2024)
Enhancing Traffic Safety with Parallel Dense Video Captioning for End-to-End Event Analysis
di: Shoman, Maged, et al.
Pubblicazione: (2024)
di: Shoman, Maged, et al.
Pubblicazione: (2024)
Multi-view Structural Convolution Network for Domain-Invariant Point Cloud Recognition of Autonomous Vehicles
di: Kim, Younggun, et al.
Pubblicazione: (2025)
di: Kim, Younggun, et al.
Pubblicazione: (2025)
Enhancing Vision-Language Models with Scene Graphs for Traffic Accident Understanding
di: Lohner, Aaron, et al.
Pubblicazione: (2024)
di: Lohner, Aaron, et al.
Pubblicazione: (2024)
AVD2: Accident Video Diffusion for Accident Video Description
di: Li, Cheng, et al.
Pubblicazione: (2025)
di: Li, Cheng, et al.
Pubblicazione: (2025)
Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety
di: Kim, Younggun, et al.
Pubblicazione: (2025)
di: Kim, Younggun, et al.
Pubblicazione: (2025)
Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models
di: Mohamud, Safaa Abdullahi Moallim, et al.
Pubblicazione: (2025)
di: Mohamud, Safaa Abdullahi Moallim, et al.
Pubblicazione: (2025)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
di: Jeon, MinJu, et al.
Pubblicazione: (2025)
di: Jeon, MinJu, et al.
Pubblicazione: (2025)
VAGNet: Vision-based Accident Anticipation with Global Features
di: Vipulananthan, Vipooshan, et al.
Pubblicazione: (2026)
di: Vipulananthan, Vipooshan, et al.
Pubblicazione: (2026)
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
di: Drago, Mauro Orazio, et al.
Pubblicazione: (2025)
di: Drago, Mauro Orazio, et al.
Pubblicazione: (2025)
Abductive Ego-View Accident Video Understanding for Safe Driving Perception
di: Fang, Jianwu, et al.
Pubblicazione: (2024)
di: Fang, Jianwu, et al.
Pubblicazione: (2024)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
di: Zou, Bo, et al.
Pubblicazione: (2024)
di: Zou, Bo, et al.
Pubblicazione: (2024)
3D Question Answering for City Scene Understanding
di: Sun, Penglei, et al.
Pubblicazione: (2024)
di: Sun, Penglei, et al.
Pubblicazione: (2024)
Question-Answering Dense Video Events
di: Qin, Hangyu, et al.
Pubblicazione: (2024)
di: Qin, Hangyu, et al.
Pubblicazione: (2024)
Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
di: Ataallah, Kirolos, et al.
Pubblicazione: (2024)
di: Ataallah, Kirolos, et al.
Pubblicazione: (2024)
SAVeD: A First-Person Social Media Video Dataset for ADAS-equipped vehicle Near-Miss and Crash Event Analyses
di: Zhai, Shaoyan, et al.
Pubblicazione: (2025)
di: Zhai, Shaoyan, et al.
Pubblicazione: (2025)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
di: Chen, Guo, et al.
Pubblicazione: (2024)
di: Chen, Guo, et al.
Pubblicazione: (2024)
Streaming Dense Video Captioning
di: Zhou, Xingyi, et al.
Pubblicazione: (2024)
di: Zhou, Xingyi, et al.
Pubblicazione: (2024)
AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering
di: Chen, Xiuyuan, et al.
Pubblicazione: (2023)
di: Chen, Xiuyuan, et al.
Pubblicazione: (2023)
Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
di: Li, Lei-lei, et al.
Pubblicazione: (2025)
di: Li, Lei-lei, et al.
Pubblicazione: (2025)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
di: Mohamed, Abdelrahman, et al.
Pubblicazione: (2025)
di: Mohamed, Abdelrahman, et al.
Pubblicazione: (2025)
ACCIDENT: A Benchmark Dataset for Vehicle Accident Detection from Traffic Surveillance Videos
di: Picek, Lukas, et al.
Pubblicazione: (2026)
di: Picek, Lukas, et al.
Pubblicazione: (2026)
Multi-Stage VLM Pipeline for Zero-Shot Traffic Accident Understanding
di: Tatematsu, Fumiya, et al.
Pubblicazione: (2026)
di: Tatematsu, Fumiya, et al.
Pubblicazione: (2026)
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
di: Lu, Fan, et al.
Pubblicazione: (2024)
di: Lu, Fan, et al.
Pubblicazione: (2024)
BikeActions: An Open Platform and Benchmark for Cyclist-Centric VRU Action Recognition
di: Buettner, Max A., et al.
Pubblicazione: (2026)
di: Buettner, Max A., et al.
Pubblicazione: (2026)
DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes
di: Al-Mohannadi, Aisha, et al.
Pubblicazione: (2026)
di: Al-Mohannadi, Aisha, et al.
Pubblicazione: (2026)
Time-Scaling State-Space Models for Dense Video Captioning
di: Piergiovanni, AJ, et al.
Pubblicazione: (2025)
di: Piergiovanni, AJ, et al.
Pubblicazione: (2025)
AccidentSim: Generating Vehicle Collision Videos with Physically Realistic Collision Trajectories from Real-World Accident Reports
di: Zhang, Xiangwen, et al.
Pubblicazione: (2025)
di: Zhang, Xiangwen, et al.
Pubblicazione: (2025)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
di: Wang, Xingrui, et al.
Pubblicazione: (2024)
di: Wang, Xingrui, et al.
Pubblicazione: (2024)
LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
di: Zhang, Hongjie, et al.
Pubblicazione: (2023)
di: Zhang, Hongjie, et al.
Pubblicazione: (2023)
EQ-TAA: Equivariant Traffic Accident Anticipation via Diffusion-Based Accident Video Synthesis
di: Fang, Jianwu, et al.
Pubblicazione: (2025)
di: Fang, Jianwu, et al.
Pubblicazione: (2025)
Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning
di: Kapuriya, Janak, et al.
Pubblicazione: (2025)
di: Kapuriya, Janak, et al.
Pubblicazione: (2025)
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
di: Kim, Minkuk, et al.
Pubblicazione: (2024)
di: Kim, Minkuk, et al.
Pubblicazione: (2024)
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
di: Kim, Minkuk, et al.
Pubblicazione: (2024)
di: Kim, Minkuk, et al.
Pubblicazione: (2024)
SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
di: Mkhallati, Hassan, et al.
Pubblicazione: (2023)
di: Mkhallati, Hassan, et al.
Pubblicazione: (2023)
Scene-Text Grounding for Text-Based Video Question Answering
di: Zhou, Sheng, et al.
Pubblicazione: (2024)
di: Zhou, Sheng, et al.
Pubblicazione: (2024)
Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering
di: Anaissi, Ali, et al.
Pubblicazione: (2025)
di: Anaissi, Ali, et al.
Pubblicazione: (2025)
When, Where, and What? A Novel Benchmark for Accident Anticipation and Localization with Large Language Models
di: Liao, Haicheng, et al.
Pubblicazione: (2024)
di: Liao, Haicheng, et al.
Pubblicazione: (2024)
HORNet: Task-Guided Frame Selection for Video Question Answering with Vision-Language Models
di: Bai, Xiangyu, et al.
Pubblicazione: (2026)
di: Bai, Xiangyu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
VRU-CIPI: Crossing Intention Prediction at Intersections for Improving Vulnerable Road Users Safety
di: Abdelrahman, Ahmed S., et al.
Pubblicazione: (2025) -
Video-to-Text Pedestrian Monitoring (VTPM): Leveraging Computer Vision and Large Language Models for Privacy-Preserve Pedestrian Activity Monitoring at Intersections
di: Abdelrahman, Ahmed S., et al.
Pubblicazione: (2024) -
Enhancing Traffic Safety with Parallel Dense Video Captioning for End-to-End Event Analysis
di: Shoman, Maged, et al.
Pubblicazione: (2024) -
Multi-view Structural Convolution Network for Domain-Invariant Point Cloud Recognition of Autonomous Vehicles
di: Kim, Younggun, et al.
Pubblicazione: (2025) -
Enhancing Vision-Language Models with Scene Graphs for Traffic Accident Understanding
di: Lohner, Aaron, et al.
Pubblicazione: (2024)