Seeing Beyond Frames: Zero-Shot Pedestrian Intention Prediction with Raw Temporal Video and Multimodal Cues
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zambare, Pallavi, Thanikella, Venkata Nikhil, Liu, Ying |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Securing Agentic AI: Threat Modeling and Risk Analysis for Network Monitoring Agentic AI System
von: Zambare, Pallavi, et al.
Veröffentlicht: (2025)
von: Zambare, Pallavi, et al.
Veröffentlicht: (2025)
NetMoniAI: An Agentic AI Framework for Network Security & Monitoring
von: Zambare, Pallavi, et al.
Veröffentlicht: (2025)
von: Zambare, Pallavi, et al.
Veröffentlicht: (2025)
Pedestrian Crossing Intention Prediction Using Multimodal Fusion Network
von: Li, Yuanzhe, et al.
Veröffentlicht: (2025)
von: Li, Yuanzhe, et al.
Veröffentlicht: (2025)
Intention-Aware Diffusion Model for Pedestrian Trajectory Prediction
von: Liu, Yu, et al.
Veröffentlicht: (2025)
von: Liu, Yu, et al.
Veröffentlicht: (2025)
Occlusion-Aware Diffusion Model for Pedestrian Intention Prediction
von: Liu, Yu, et al.
Veröffentlicht: (2025)
von: Liu, Yu, et al.
Veröffentlicht: (2025)
Investigation of Frame Differences as Motion Cues for Video Object Segmentation
von: Kawamura, Sota, et al.
Veröffentlicht: (2025)
von: Kawamura, Sota, et al.
Veröffentlicht: (2025)
Multi-Context Fusion Transformer for Pedestrian Crossing Intention Prediction in Urban Environments
von: Li, Yuanzhe, et al.
Veröffentlicht: (2025)
von: Li, Yuanzhe, et al.
Veröffentlicht: (2025)
ESIA: An Energy-Based Spatiotemporal Interaction-Aware Framework for Pedestrian Intention Prediction
von: Wu, Yanping, et al.
Veröffentlicht: (2026)
von: Wu, Yanping, et al.
Veröffentlicht: (2026)
PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction
von: Mishra, Naman, et al.
Veröffentlicht: (2026)
von: Mishra, Naman, et al.
Veröffentlicht: (2026)
Zero-Shot Temporal Interaction Localization for Egocentric Videos
von: Zhang, Erhang, et al.
Veröffentlicht: (2025)
von: Zhang, Erhang, et al.
Veröffentlicht: (2025)
VTG-GPT: Tuning-Free Zero-Shot Video Temporal Grounding with GPT
von: Xu, Yifang, et al.
Veröffentlicht: (2024)
von: Xu, Yifang, et al.
Veröffentlicht: (2024)
Zero-TIG: Temporal Consistency-Aware Zero-Shot Illumination-Guided Low-light Video Enhancement
von: Li, Yini, et al.
Veröffentlicht: (2025)
von: Li, Yini, et al.
Veröffentlicht: (2025)
Temporal-contextual Event Learning for Pedestrian Crossing Intent Prediction
von: Liang, Hongbin, et al.
Veröffentlicht: (2025)
von: Liang, Hongbin, et al.
Veröffentlicht: (2025)
Beyond Frequency: Seeing Subtle Cues Through the Lens of Spatial Decomposition for Fine-Grained Visual Classification
von: Xu, Qin, et al.
Veröffentlicht: (2025)
von: Xu, Qin, et al.
Veröffentlicht: (2025)
Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification
von: Wang, Xiao, et al.
Veröffentlicht: (2026)
von: Wang, Xiao, et al.
Veröffentlicht: (2026)
ART: Adaptive Relational Transformer for Pedestrian Trajectory Prediction with Temporal-Aware Relations
von: Li, Ruochen, et al.
Veröffentlicht: (2026)
von: Li, Ruochen, et al.
Veröffentlicht: (2026)
Unified Spatial-Temporal Edge-Enhanced Graph Networks for Pedestrian Trajectory Prediction
von: Li, Ruochen, et al.
Veröffentlicht: (2025)
von: Li, Ruochen, et al.
Veröffentlicht: (2025)
On the Reliability of Cue Conflict and Beyond
von: Kim, Pum Jun, et al.
Veröffentlicht: (2026)
von: Kim, Pum Jun, et al.
Veröffentlicht: (2026)
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
von: Zhang, Peng, et al.
Veröffentlicht: (2026)
von: Zhang, Peng, et al.
Veröffentlicht: (2026)
VIT-Ped: Visionary Intention Transformer for Pedestrian Behavior Analysis
von: Elkammar, Aly R., et al.
Veröffentlicht: (2026)
von: Elkammar, Aly R., et al.
Veröffentlicht: (2026)
Synthetic Data Generation Framework, Dataset, and Efficient Deep Model for Pedestrian Intention Prediction
von: Riaz, Muhammad Naveed, et al.
Veröffentlicht: (2024)
von: Riaz, Muhammad Naveed, et al.
Veröffentlicht: (2024)
GVDepth: Zero-Shot Monocular Depth Estimation for Ground Vehicles based on Probabilistic Cue Fusion
von: Koledić, Karlo, et al.
Veröffentlicht: (2024)
von: Koledić, Karlo, et al.
Veröffentlicht: (2024)
Spatio-Temporal Context Prompting for Zero-Shot Action Detection
von: Huang, Wei-Jhe, et al.
Veröffentlicht: (2024)
von: Huang, Wei-Jhe, et al.
Veröffentlicht: (2024)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
von: Ji, Longbin, et al.
Veröffentlicht: (2026)
von: Ji, Longbin, et al.
Veröffentlicht: (2026)
CaTFormer: Causal Temporal Transformer with Dynamic Contextual Fusion for Driving Intention Prediction
von: Wang, Sirui, et al.
Veröffentlicht: (2025)
von: Wang, Sirui, et al.
Veröffentlicht: (2025)
Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
von: Guo, Xingang, et al.
Veröffentlicht: (2025)
von: Guo, Xingang, et al.
Veröffentlicht: (2025)
Benefits of Feature Extraction and Temporal Sequence Analysis for Video Frame Prediction: An Evaluation of Hybrid Deep Learning Models
von: Velázquez, Jose M. Sánchez, et al.
Veröffentlicht: (2025)
von: Velázquez, Jose M. Sánchez, et al.
Veröffentlicht: (2025)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
Beyond Pedestrians: Caption-Guided CLIP Framework for High-Difficulty Video-based Person Re-Identification
von: Hamano, Shogo, et al.
Veröffentlicht: (2026)
von: Hamano, Shogo, et al.
Veröffentlicht: (2026)
VideoPoet: A Large Language Model for Zero-Shot Video Generation
von: Kondratyuk, Dan, et al.
Veröffentlicht: (2023)
von: Kondratyuk, Dan, et al.
Veröffentlicht: (2023)
See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection
von: Wu, Zhiheng, et al.
Veröffentlicht: (2026)
von: Wu, Zhiheng, et al.
Veröffentlicht: (2026)
No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection
von: Dai, Zunkai, et al.
Veröffentlicht: (2026)
von: Dai, Zunkai, et al.
Veröffentlicht: (2026)
Zero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing
von: Su, Tongtong, et al.
Veröffentlicht: (2025)
von: Su, Tongtong, et al.
Veröffentlicht: (2025)
DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation
von: Hong, Susung, et al.
Veröffentlicht: (2023)
von: Hong, Susung, et al.
Veröffentlicht: (2023)
Pedestrian Intention Prediction via Vision-Language Foundation Models
von: Azarmi, Mohsen, et al.
Veröffentlicht: (2025)
von: Azarmi, Mohsen, et al.
Veröffentlicht: (2025)
Bayesian Modeling of Zero-Shot Classifications for Urban Flood Detection
von: Franchi, Matt, et al.
Veröffentlicht: (2025)
von: Franchi, Matt, et al.
Veröffentlicht: (2025)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
von: Kim, Youngmin, et al.
Veröffentlicht: (2025)
von: Kim, Youngmin, et al.
Veröffentlicht: (2025)
Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC Videos
von: Saini, Shreshth, et al.
Veröffentlicht: (2026)
von: Saini, Shreshth, et al.
Veröffentlicht: (2026)
GVCC: Zero-Shot Video Compression via Codebook-Driven Stochastic Rectified Flow
von: Zeng, Ziyue, et al.
Veröffentlicht: (2026)
von: Zeng, Ziyue, et al.
Veröffentlicht: (2026)
Context-Aware Pseudo-Label Scoring for Zero-Shot Video Summarization
von: Wu, Yuanli, et al.
Veröffentlicht: (2025)
von: Wu, Yuanli, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Securing Agentic AI: Threat Modeling and Risk Analysis for Network Monitoring Agentic AI System
von: Zambare, Pallavi, et al.
Veröffentlicht: (2025) -
NetMoniAI: An Agentic AI Framework for Network Security & Monitoring
von: Zambare, Pallavi, et al.
Veröffentlicht: (2025) -
Pedestrian Crossing Intention Prediction Using Multimodal Fusion Network
von: Li, Yuanzhe, et al.
Veröffentlicht: (2025) -
Intention-Aware Diffusion Model for Pedestrian Trajectory Prediction
von: Liu, Yu, et al.
Veröffentlicht: (2025) -
Occlusion-Aware Diffusion Model for Pedestrian Intention Prediction
von: Liu, Yu, et al.
Veröffentlicht: (2025)