GPT-4V Takes the Wheel: Promises and Challenges for Pedestrian Behavior Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Jia, Jiang, Peng, Gautam, Alvika, Saripalli, Srikanth |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AnyTraverse: An off-road traversability framework with VLM and human operator in the loop
by: Sahu, Sattwik, et al.
Published: (2025)
by: Sahu, Sattwik, et al.
Published: (2025)
3DGS-ReLoc: 3D Gaussian Splatting for Map Representation and Visual ReLocalization
by: Jiang, Peng, et al.
Published: (2024)
by: Jiang, Peng, et al.
Published: (2024)
Off-Road LiDAR Intensity Based Semantic Segmentation
by: Viswanath, Kasi, et al.
Published: (2024)
by: Viswanath, Kasi, et al.
Published: (2024)
Reflectivity Is All You Need!: Advancing LiDAR Semantic Segmentation
by: Viswanath, Kasi, et al.
Published: (2024)
by: Viswanath, Kasi, et al.
Published: (2024)
VIT-Ped: Visionary Intention Transformer for Pedestrian Behavior Analysis
by: Elkammar, Aly R., et al.
Published: (2026)
by: Elkammar, Aly R., et al.
Published: (2026)
Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V
by: Zhi, Peiyuan, et al.
Published: (2024)
by: Zhi, Peiyuan, et al.
Published: (2024)
OFFSEG: A Semantic Segmentation Framework For Off-Road Driving
by: Viswanath, Kasi, et al.
Published: (2021)
by: Viswanath, Kasi, et al.
Published: (2021)
Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
by: Gao, Haoxiang, et al.
Published: (2025)
by: Gao, Haoxiang, et al.
Published: (2025)
G-PECNet: Towards a Generalizable Pedestrian Trajectory Prediction System
by: Garg, Aryan, et al.
Published: (2022)
by: Garg, Aryan, et al.
Published: (2022)
Learning the Pedestrian-Vehicle Interaction for Pedestrian Trajectory Prediction
by: Zhang, Chi, et al.
Published: (2022)
by: Zhang, Chi, et al.
Published: (2022)
DriveGPT: Scaling Autoregressive Behavior Models for Driving
by: Huang, Xin, et al.
Published: (2024)
by: Huang, Xin, et al.
Published: (2024)
Controllable Pedestrian Video Editing for Multi-View Driving Scenarios via Motion Sequence
by: Fu, Danzhen, et al.
Published: (2025)
by: Fu, Danzhen, et al.
Published: (2025)
A Spatio-temporal Graph Network Allowing Incomplete Trajectory Input for Pedestrian Trajectory Prediction
by: Long, Juncen, et al.
Published: (2025)
by: Long, Juncen, et al.
Published: (2025)
Research on Reliable and Safe Occupancy Grid Prediction in Underground Parking Lots
by: Luo, JiaQi
Published: (2024)
by: Luo, JiaQi
Published: (2024)
Automated Lane Change Behavior Prediction and Environmental Perception Based on SLAM Technology
by: Lei, Han, et al.
Published: (2024)
by: Lei, Han, et al.
Published: (2024)
mmWave Radar-Based Non-Line-of-Sight Pedestrian Localization at T-Junctions Utilizing Road Layout Extraction via Camera
by: Park, Byeonggyu, et al.
Published: (2025)
by: Park, Byeonggyu, et al.
Published: (2025)
Pedestrian Intention Prediction via Vision-Language Foundation Models
by: Azarmi, Mohsen, et al.
Published: (2025)
by: Azarmi, Mohsen, et al.
Published: (2025)
CSAOT: Cooperative Multi-Agent System for Active Object Tracking
by: Nguyen, Hy, et al.
Published: (2025)
by: Nguyen, Hy, et al.
Published: (2025)
From Scene to Object: Text-Guided Dual-Gaze Prediction
by: Ke, Zehong, et al.
Published: (2026)
by: Ke, Zehong, et al.
Published: (2026)
Efficient Driving Behavior Narration and Reasoning on Edge Device Using Large Language Models
by: Huang, Yizhou, et al.
Published: (2024)
by: Huang, Yizhou, et al.
Published: (2024)
4th Workshop on Maritime Computer Vision (MaCVi): Challenge Overview
by: Kiefer, Benjamin, et al.
Published: (2026)
by: Kiefer, Benjamin, et al.
Published: (2026)
V-HOP: Visuo-Haptic 6D Object Pose Tracking
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
Sparse Prototype Network for Explainable Pedestrian Behavior Prediction
by: Feng, Yan, et al.
Published: (2024)
by: Feng, Yan, et al.
Published: (2024)
SSL-Interactions: Pretext Tasks for Interactive Trajectory Prediction
by: Bhattacharyya, Prarthana, et al.
Published: (2024)
by: Bhattacharyya, Prarthana, et al.
Published: (2024)
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation
by: Chen, Jiaqi, et al.
Published: (2024)
by: Chen, Jiaqi, et al.
Published: (2024)
The Safety Challenge of World Models for Embodied AI Agents: A Review
by: Baraldi, Lorenzo, et al.
Published: (2025)
by: Baraldi, Lorenzo, et al.
Published: (2025)
Characterizing Structured versus Unstructured Environments based on Pedestrians' and Vehicles' Motion Trajectories
by: Golchoubian, Mahsa, et al.
Published: (2025)
by: Golchoubian, Mahsa, et al.
Published: (2025)
SpatialPoint: Spatial-aware Point Prediction for Embodied Localization
by: Zhu, Qiming, et al.
Published: (2026)
by: Zhu, Qiming, et al.
Published: (2026)
V2X-QA: A Comprehensive Reasoning Dataset and Benchmark for Multimodal Large Language Models in Autonomous Driving Across Ego, Infrastructure, and Cooperative Views
by: You, Junwei, et al.
Published: (2026)
by: You, Junwei, et al.
Published: (2026)
Multi-modal Situated Reasoning in 3D Scenes
by: Linghu, Xiongkun, et al.
Published: (2024)
by: Linghu, Xiongkun, et al.
Published: (2024)
Feature Importance in Pedestrian Intention Prediction: A Context-Aware Review
by: Azarmi, Mohsen, et al.
Published: (2024)
by: Azarmi, Mohsen, et al.
Published: (2024)
Advances and Innovations in the Multi-Agent Robotic System (MARS) Challenge
by: Kang, Li, et al.
Published: (2026)
by: Kang, Li, et al.
Published: (2026)
WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models
by: Chen, Hongjin, et al.
Published: (2026)
by: Chen, Hongjin, et al.
Published: (2026)
Behavioral Cloning Models Reality Check for Autonomous Driving
by: Yildirim, Mustafa, et al.
Published: (2024)
by: Yildirim, Mustafa, et al.
Published: (2024)
SKT: Integrating State-Aware Keypoint Trajectories with Vision-Language Models for Robotic Garment Manipulation
by: Li, Xin, et al.
Published: (2024)
by: Li, Xin, et al.
Published: (2024)
TFusionOcc: T-Primitive Based Object-Centric Multi-Sensor Fusion Framework for 3D Occupancy Prediction
by: Ming, Zhenxing, et al.
Published: (2026)
by: Ming, Zhenxing, et al.
Published: (2026)
No Pedestrian Left Behind: Real-Time Detection and Tracking of Vulnerable Road Users for Adaptive Traffic Signal Control
by: Aly, Anas Gamal, et al.
Published: (2026)
by: Aly, Anas Gamal, et al.
Published: (2026)
RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis
by: Mu, Yao, et al.
Published: (2024)
by: Mu, Yao, et al.
Published: (2024)
EC-Diffuser: Multi-Object Manipulation via Entity-Centric Behavior Generation
by: Qi, Carl, et al.
Published: (2024)
by: Qi, Carl, et al.
Published: (2024)
Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
by: Qi, Xiuxiu, et al.
Published: (2025)
by: Qi, Xiuxiu, et al.
Published: (2025)
Similar Items
-
AnyTraverse: An off-road traversability framework with VLM and human operator in the loop
by: Sahu, Sattwik, et al.
Published: (2025) -
3DGS-ReLoc: 3D Gaussian Splatting for Map Representation and Visual ReLocalization
by: Jiang, Peng, et al.
Published: (2024) -
Off-Road LiDAR Intensity Based Semantic Segmentation
by: Viswanath, Kasi, et al.
Published: (2024) -
Reflectivity Is All You Need!: Advancing LiDAR Semantic Segmentation
by: Viswanath, Kasi, et al.
Published: (2024) -
VIT-Ped: Visionary Intention Transformer for Pedestrian Behavior Analysis
by: Elkammar, Aly R., et al.
Published: (2026)