Multi-Stage VLM Pipeline for Zero-Shot Traffic Accident Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Tatematsu, Fumiya, Takahashi, Fumihiko |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Modular Zero-Shot Pipeline for Accident Detection, Localization, and Classification in Traffic Surveillance Video
by: Thakur, Amey, et al.
Published: (2026)
by: Thakur, Amey, et al.
Published: (2026)
ContextVLM: Zero-Shot and Few-Shot Context Understanding for Autonomous Driving using Vision Language Models
by: Sural, Shounak, et al.
Published: (2024)
by: Sural, Shounak, et al.
Published: (2024)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
StoryTailor:A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual Narratives
by: Hu, Jinghao, et al.
Published: (2026)
by: Hu, Jinghao, et al.
Published: (2026)
SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding
by: Lin, Jiawen, et al.
Published: (2025)
by: Lin, Jiawen, et al.
Published: (2025)
HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation
by: Lei, Qinqian, et al.
Published: (2025)
by: Lei, Qinqian, et al.
Published: (2025)
EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI Detection
by: Lei, Qinqian, et al.
Published: (2024)
by: Lei, Qinqian, et al.
Published: (2024)
Real-time Traffic Accident Anticipation with Feature Reuse
by: Song, Inpyo, et al.
Published: (2025)
by: Song, Inpyo, et al.
Published: (2025)
Enhancing Vision-Language Models with Scene Graphs for Traffic Accident Understanding
by: Lohner, Aaron, et al.
Published: (2024)
by: Lohner, Aaron, et al.
Published: (2024)
Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
by: Li, Lei-lei, et al.
Published: (2025)
by: Li, Lei-lei, et al.
Published: (2025)
Leveraging VLM-Based Pipelines to Annotate 3D Objects
by: Kabra, Rishabh, et al.
Published: (2023)
by: Kabra, Rishabh, et al.
Published: (2023)
Zero-Shot Long-Form Video Understanding through Screenplay
by: Wu, Yongliang, et al.
Published: (2024)
by: Wu, Yongliang, et al.
Published: (2024)
Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion
by: Klepachevskyi, Dmytro, et al.
Published: (2026)
by: Klepachevskyi, Dmytro, et al.
Published: (2026)
MsFIN: Multi-scale Feature Interaction Network for Traffic Accident Anticipation
by: Wu, Tongshuai, et al.
Published: (2025)
by: Wu, Tongshuai, et al.
Published: (2025)
Two-Pass Zero-Shot Temporal-Spatial Grounding of Rare Traffic Events in Surveillance Video
by: Huang, Jiantang
Published: (2026)
by: Huang, Jiantang
Published: (2026)
Zero-Shot Multi-Animal Tracking in the Wild
by: Meier, Jan Frederik, et al.
Published: (2025)
by: Meier, Jan Frederik, et al.
Published: (2025)
MVSAnywhere: Zero-Shot Multi-View Stereo
by: Izquierdo, Sergio, et al.
Published: (2025)
by: Izquierdo, Sergio, et al.
Published: (2025)
Zero-Shot Multi-Object Scene Completion
by: Iwase, Shun, et al.
Published: (2024)
by: Iwase, Shun, et al.
Published: (2024)
Safety-Critical Learning for Long-Tail Events: The TUM Traffic Accident Dataset
by: Zimmer, Walter, et al.
Published: (2025)
by: Zimmer, Walter, et al.
Published: (2025)
SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding
by: Sheng, Zihao, et al.
Published: (2025)
by: Sheng, Zihao, et al.
Published: (2025)
T2ICount: Enhancing Cross-modal Understanding for Zero-Shot Counting
by: Qian, Yifei, et al.
Published: (2025)
by: Qian, Yifei, et al.
Published: (2025)
A Recipe for Improving Remote Sensing VLM Zero Shot Generalization
by: Barzilai, Aviad, et al.
Published: (2025)
by: Barzilai, Aviad, et al.
Published: (2025)
MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval
by: Tu, Rong-Cheng, et al.
Published: (2025)
by: Tu, Rong-Cheng, et al.
Published: (2025)
VLM-in-the-Loop: A Plug-In Quality Assurance Module for ECG Digitization Pipelines
by: Li, Jiachen, et al.
Published: (2026)
by: Li, Jiachen, et al.
Published: (2026)
AccidentGPT: Large Multi-Modal Foundation Model for Traffic Accident Analysis
by: Wu, Kebin, et al.
Published: (2024)
by: Wu, Kebin, et al.
Published: (2024)
MLFM: Multi-Layered Feature Maps for Richer Language Understanding in Zero-Shot Semantic Navigation
by: Raychaudhuri, Sonia, et al.
Published: (2025)
by: Raychaudhuri, Sonia, et al.
Published: (2025)
Benchmarking and Enhancing VLM for Compressed Image Understanding
by: Zhang, Zifu, et al.
Published: (2025)
by: Zhang, Zifu, et al.
Published: (2025)
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
by: Kim, Younggun, et al.
Published: (2025)
by: Kim, Younggun, et al.
Published: (2025)
Vision-Language Integration for Zero-Shot Scene Understanding in Real-World Environments
by: Rajiv, Manjunath Prasad Holenarasipura, et al.
Published: (2025)
by: Rajiv, Manjunath Prasad Holenarasipura, et al.
Published: (2025)
TinyVLM: Zero-Shot Object Detection on Microcontrollers via Vision-Language Distillation with Matryoshka Embeddings
by: Wilson, Bibin
Published: (2026)
by: Wilson, Bibin
Published: (2026)
A Multi-Stage Optimization Pipeline for Bethesda Cell Detection in Pap Smear Cytology
by: Amster, Martin, et al.
Published: (2026)
by: Amster, Martin, et al.
Published: (2026)
Will It Zero-Shot?: Predicting Zero-Shot Classification Performance For Arbitrary Queries
by: Robbins, Kevin, et al.
Published: (2026)
by: Robbins, Kevin, et al.
Published: (2026)
Structure-Preserving Zero-Shot Image Editing via Stage-Wise Latent Injection in Diffusion Models
by: Jeong, Dasol, et al.
Published: (2025)
by: Jeong, Dasol, et al.
Published: (2025)
Integrating Generative Adversarial Networks and Convolutional Neural Networks for Enhanced Traffic Accidents Detection and Analysis
by: Xi, Zhenghao, et al.
Published: (2025)
by: Xi, Zhenghao, et al.
Published: (2025)
Multi-Granularity Mutual Refinement Network for Zero-Shot Learning
by: Wang, Ning, et al.
Published: (2025)
by: Wang, Ning, et al.
Published: (2025)
MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM
by: Chen, Tao, et al.
Published: (2025)
by: Chen, Tao, et al.
Published: (2025)
Focus-Consistent Multi-Level Aggregation for Compositional Zero-Shot Learning
by: Dai, Fengyuan, et al.
Published: (2024)
by: Dai, Fengyuan, et al.
Published: (2024)
Multi-Scale Memory Comparison for Zero-/Few-Shot Anomaly Detection
by: Huang, Chaoqin, et al.
Published: (2023)
by: Huang, Chaoqin, et al.
Published: (2023)
MADS: Multi-Attribute Document Supervision for Zero-Shot Image Classification
by: Qu, Xiangyan, et al.
Published: (2025)
by: Qu, Xiangyan, et al.
Published: (2025)
Multi-Teacher Multi-Objective Meta-Learning for Zero-Shot Hyperspectral Band Selection
by: Feng, Jie, et al.
Published: (2024)
by: Feng, Jie, et al.
Published: (2024)
Similar Items
-
A Modular Zero-Shot Pipeline for Accident Detection, Localization, and Classification in Traffic Surveillance Video
by: Thakur, Amey, et al.
Published: (2026) -
ContextVLM: Zero-Shot and Few-Shot Context Understanding for Autonomous Driving using Vision Language Models
by: Sural, Shounak, et al.
Published: (2024) -
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024) -
StoryTailor:A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual Narratives
by: Hu, Jinghao, et al.
Published: (2026) -
SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding
by: Lin, Jiawen, et al.
Published: (2025)