Enhancing Traffic Safety with Parallel Dense Video Captioning for End-to-End Event Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Shoman, Maged, Wang, Dongdong, Aboah, Armstrong, Abdel-Aty, Mohamed |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
by: Kim, Younggun, et al.
Published: (2025)
by: Kim, Younggun, et al.
Published: (2025)
Low-Light Image Enhancement Framework for Improved Object Detection in Fisheye Lens Datasets
by: Tran, Dai Quoc, et al.
Published: (2024)
by: Tran, Dai Quoc, et al.
Published: (2024)
End-to-End Dense Video Grounding via Parallel Regression
by: Shi, Fengyuan, et al.
Published: (2021)
by: Shi, Fengyuan, et al.
Published: (2021)
Video-to-Text Pedestrian Monitoring (VTPM): Leveraging Computer Vision and Large Language Models for Privacy-Preserve Pedestrian Activity Monitoring at Intersections
by: Abdelrahman, Ahmed S., et al.
Published: (2024)
by: Abdelrahman, Ahmed S., et al.
Published: (2024)
PaveCap: The First Multimodal Framework for Comprehensive Pavement Condition Assessment with Dense Captioning and PCI Estimation
by: Kyem, Blessing Agyei, et al.
Published: (2024)
by: Kyem, Blessing Agyei, et al.
Published: (2024)
SAVeD: A First-Person Social Media Video Dataset for ADAS-equipped vehicle Near-Miss and Crash Event Analyses
by: Zhai, Shaoyan, et al.
Published: (2025)
by: Zhai, Shaoyan, et al.
Published: (2025)
A Contextual Analysis of Driver-Facing and Dual-View Video Inputs for Distraction Detection in Naturalistic Driving Environments
by: Dontoh, Anthony, et al.
Published: (2025)
by: Dontoh, Anthony, et al.
Published: (2025)
Task-Specific Dual-Model Framework for Comprehensive Traffic Safety Video Description and Analysis
by: Kyem, Blessing Agyei, et al.
Published: (2025)
by: Kyem, Blessing Agyei, et al.
Published: (2025)
A Road-Conditioned Traffic Movie Prediction Network with Spatiotemporal and Structure-Consistent Learning
by: Asamoah, Joshua Kofi, et al.
Published: (2026)
by: Asamoah, Joshua Kofi, et al.
Published: (2026)
Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-learning
by: Xie, Zhuyang, et al.
Published: (2024)
by: Xie, Zhuyang, et al.
Published: (2024)
An Integrated Causal Inference Framework for Traffic Safety Modeling with Semantic Street-View Visual Features
by: Sun, Lishan, et al.
Published: (2026)
by: Sun, Lishan, et al.
Published: (2026)
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024)
by: Zhou, Xingyi, et al.
Published: (2024)
Exploiting Multiple Sequence Lengths in Fast End to End Training for Image Captioning
by: Hu, Jia Cheng, et al.
Published: (2022)
by: Hu, Jia Cheng, et al.
Published: (2022)
An Improved ResNet50 Model for Predicting Pavement Condition Index (PCI) Directly from Pavement Images
by: Danyo, Andrews, et al.
Published: (2025)
by: Danyo, Andrews, et al.
Published: (2025)
VRU-CIPI: Crossing Intention Prediction at Intersections for Improving Vulnerable Road Users Safety
by: Abdelrahman, Ahmed S., et al.
Published: (2025)
by: Abdelrahman, Ahmed S., et al.
Published: (2025)
Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving
by: Sheng, Zihao, et al.
Published: (2026)
by: Sheng, Zihao, et al.
Published: (2026)
SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
by: Mkhallati, Hassan, et al.
Published: (2023)
by: Mkhallati, Hassan, et al.
Published: (2023)
LMVC: An End-to-End Learned Multiview Video Coding Framework
by: Sheng, Xihua, et al.
Published: (2025)
by: Sheng, Xihua, et al.
Published: (2025)
RT-DETRv3: Real-time End-to-End Object Detection with Hierarchical Dense Positive Supervision
by: Wang, Shuo, et al.
Published: (2024)
by: Wang, Shuo, et al.
Published: (2024)
End-to-End Facial Expression Detection in Long Videos
by: Fang, Yini, et al.
Published: (2025)
by: Fang, Yini, et al.
Published: (2025)
EVC-MF: End-to-end Video Captioning Network with Multi-scale Features
by: Niu, Tian-Zi, et al.
Published: (2024)
by: Niu, Tian-Zi, et al.
Published: (2024)
DLAFormer: An End-to-End Transformer For Document Layout Analysis
by: Wang, Jiawei, et al.
Published: (2024)
by: Wang, Jiawei, et al.
Published: (2024)
Technical Report for Soccernet 2023 -- Dense Video Captioning
by: Ruan, Zheng, et al.
Published: (2024)
by: Ruan, Zheng, et al.
Published: (2024)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
by: Jeon, MinJu, et al.
Published: (2025)
by: Jeon, MinJu, et al.
Published: (2025)
Dense Video Captioning Using Unsupervised Semantic Information
by: Estevam, Valter, et al.
Published: (2021)
by: Estevam, Valter, et al.
Published: (2021)
Dense Video Object Captioning from Disjoint Supervision
by: Zhou, Xingyi, et al.
Published: (2023)
by: Zhou, Xingyi, et al.
Published: (2023)
FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM
by: Wu, Yuchen, et al.
Published: (2025)
by: Wu, Yuchen, et al.
Published: (2025)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
by: Shi, Yudi, et al.
Published: (2026)
by: Shi, Yudi, et al.
Published: (2026)
On the Pitfalls of Batch Normalization for End-to-End Video Learning: A Study on Surgical Workflow Analysis
by: Rivoir, Dominik, et al.
Published: (2022)
by: Rivoir, Dominik, et al.
Published: (2022)
Context-CrackNet: A Context-Aware Framework for Precise Segmentation of Tiny Cracks in Pavement images
by: Kyem, Blessing Agyei, et al.
Published: (2025)
by: Kyem, Blessing Agyei, et al.
Published: (2025)
End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer
by: Yu, Yonghui, et al.
Published: (2025)
by: Yu, Yonghui, et al.
Published: (2025)
3D Object Detection and High-Resolution Traffic Parameters Extraction Using Low-Resolution LiDAR Data
by: Zhang, Linlin, et al.
Published: (2024)
by: Zhang, Linlin, et al.
Published: (2024)
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning
by: Li, Yizhe, et al.
Published: (2025)
by: Li, Yizhe, et al.
Published: (2025)
Enhancing End-to-End Autonomous Driving with Latent World Model
by: Li, Yingyan, et al.
Published: (2024)
by: Li, Yingyan, et al.
Published: (2024)
An End-to-End Framework for Video Multi-Person Pose Estimation
by: Wei, Zhihong
Published: (2025)
by: Wei, Zhihong
Published: (2025)
Enhancing Maritime Situational Awareness through End-to-End Onboard Raw Data Analysis
by: Del Prete, Roberto, et al.
Published: (2024)
by: Del Prete, Roberto, et al.
Published: (2024)
DriveSafer: End-to-End Autonomous Driving with Safety Guidance
by: Sural, Shounak, et al.
Published: (2026)
by: Sural, Shounak, et al.
Published: (2026)
MoCha:End-to-End Video Character Replacement without Structural Guidance
by: Xu, Zhengbo, et al.
Published: (2026)
by: Xu, Zhengbo, et al.
Published: (2026)
Bias at the End of the Score
by: Magid, Salma Abdel, et al.
Published: (2026)
by: Magid, Salma Abdel, et al.
Published: (2026)
Similar Items
-
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
by: Kim, Younggun, et al.
Published: (2025) -
Low-Light Image Enhancement Framework for Improved Object Detection in Fisheye Lens Datasets
by: Tran, Dai Quoc, et al.
Published: (2024) -
End-to-End Dense Video Grounding via Parallel Regression
by: Shi, Fengyuan, et al.
Published: (2021) -
Video-to-Text Pedestrian Monitoring (VTPM): Leveraging Computer Vision and Large Language Models for Privacy-Preserve Pedestrian Activity Monitoring at Intersections
by: Abdelrahman, Ahmed S., et al.
Published: (2024) -
PaveCap: The First Multimodal Framework for Comprehensive Pavement Condition Assessment with Dense Captioning and PCI Estimation
by: Kyem, Blessing Agyei, et al.
Published: (2024)