Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach
Fuente:
arXiv
Saved in:
| Main Authors: | Khindkar, Vaishnavi, Balasubramanian, Vineeth, Arora, Chetan, Subramanian, Anbumani, Jawahar, C. V. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media
by: M, Megha Mariam K., et al.
Published: (2026)
by: M, Megha Mariam K., et al.
Published: (2026)
Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning
by: Raj, Ankita, et al.
Published: (2025)
by: Raj, Ankita, et al.
Published: (2025)
Can Better Text Semantics in Prompt Tuning Improve VLM Generalization?
by: Kuchibhotla, Hari Chandana, et al.
Published: (2024)
by: Kuchibhotla, Hari Chandana, et al.
Published: (2024)
PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction
by: Mishra, Naman, et al.
Published: (2026)
by: Mishra, Naman, et al.
Published: (2026)
Source-Free Domain Adaptation by Optimizing Batch-Wise Cosine Similarity
by: Pathak, Harsharaj, et al.
Published: (2026)
by: Pathak, Harsharaj, et al.
Published: (2026)
Pedestrian Crossing Intent Prediction via Psychological Features and Transformer Fusion
by: Ashayer, Sima, et al.
Published: (2026)
by: Ashayer, Sima, et al.
Published: (2026)
Temporal-contextual Event Learning for Pedestrian Crossing Intent Prediction
by: Liang, Hongbin, et al.
Published: (2025)
by: Liang, Hongbin, et al.
Published: (2025)
Pedestrian Intention and Trajectory Prediction in Unstructured Traffic Using IDD-PeD
by: Bokkasam, Ruthvik, et al.
Published: (2025)
by: Bokkasam, Ruthvik, et al.
Published: (2025)
Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
ACIT: Attention-Guided Cross-Modal Interaction Transformer for Pedestrian Crossing Intention Prediction
by: Li, Yuanzhe, et al.
Published: (2025)
by: Li, Yuanzhe, et al.
Published: (2025)
Generative Adversarial Patches for Physical Attacks on Cross-Modal Pedestrian Re-Identification
by: Su, Yue, et al.
Published: (2024)
by: Su, Yue, et al.
Published: (2024)
C2FDrone: Coarse-to-Fine Drone-to-Drone Detection using Vision Transformer Networks
by: Rebbapragada, Sairam VC, et al.
Published: (2024)
by: Rebbapragada, Sairam VC, et al.
Published: (2024)
Robust Pedestrian Detection with Uncertain Modality
by: Bie, Qian, et al.
Published: (2026)
by: Bie, Qian, et al.
Published: (2026)
Can Generative Video Models Help Pose Estimation?
by: Cai, Ruojin, et al.
Published: (2024)
by: Cai, Ruojin, et al.
Published: (2024)
LQ-Adapter: ViT-Adapter with Learnable Queries for Gallbladder Cancer Detection from Ultrasound Image
by: Madan, Chetan, et al.
Published: (2024)
by: Madan, Chetan, et al.
Published: (2024)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
by: Mehrotra, Sarthak, et al.
Published: (2025)
by: Mehrotra, Sarthak, et al.
Published: (2025)
$\oslash$ Source Models Leak What They Shouldn't $\nrightarrow$: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimization
by: Devalapally, Arnav, et al.
Published: (2026)
by: Devalapally, Arnav, et al.
Published: (2026)
Precise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class Imbalance
by: Santra, Sanchayan, et al.
Published: (2025)
by: Santra, Sanchayan, et al.
Published: (2025)
ECoDepth: Effective Conditioning of Diffusion Models for Monocular Depth Estimation
by: Patni, Suraj, et al.
Published: (2024)
by: Patni, Suraj, et al.
Published: (2024)
Context-aware Multi-task Learning for Pedestrian Intent and Trajectory Prediction
by: Munir, Farzeen, et al.
Published: (2024)
by: Munir, Farzeen, et al.
Published: (2024)
IndicSTR12: A Dataset for Indic Scene Text Recognition
by: Lunia, Harsh, et al.
Published: (2024)
by: Lunia, Harsh, et al.
Published: (2024)
Advancing Question Answering on Handwritten Documents: A State-of-the-Art Recognition-Based Model for HW-SQuAD
by: Pal, Aniket, et al.
Published: (2024)
by: Pal, Aniket, et al.
Published: (2024)
Attend to what I say: Highlighting relevant content on slides
by: M, Megha Mariam K, et al.
Published: (2026)
by: M, Megha Mariam K, et al.
Published: (2026)
Source-free Video Domain Adaptation by Learning from Noisy Labels
by: Dasgupta, Avijit, et al.
Published: (2023)
by: Dasgupta, Avijit, et al.
Published: (2023)
LogicCBMs: Logic-Enhanced Concept-Based Learning
by: Vemuri, Deepika SN, et al.
Published: (2025)
by: Vemuri, Deepika SN, et al.
Published: (2025)
Feature Space Perturbation: A Panacea to Enhanced Transferability Estimation
by: Khoba, Prafful Kumar, et al.
Published: (2025)
by: Khoba, Prafful Kumar, et al.
Published: (2025)
Can Large Pretrained Depth Estimation Models Help With Image Dehazing?
by: Zhang, Hongfei, et al.
Published: (2025)
by: Zhang, Hongfei, et al.
Published: (2025)
Pedestrian Attribute Recognition via Hierarchical Cross-Modality HyperGraph Learning
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Can Image-To-Video Models Simulate Pedestrian Dynamics?
by: Appelle, Aaron, et al.
Published: (2025)
by: Appelle, Aaron, et al.
Published: (2025)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
by: Park, Simon, et al.
Published: (2025)
by: Park, Simon, et al.
Published: (2025)
Open-Set Object Detection By Aligning Known Class Representations
by: Sarkar, Hiran, et al.
Published: (2024)
by: Sarkar, Hiran, et al.
Published: (2024)
On Evaluation of Vision Datasets and Models using Human Competency Frameworks
by: Ramachandran, Rahul, et al.
Published: (2024)
by: Ramachandran, Rahul, et al.
Published: (2024)
Prompt2LVideos: Exploring Prompts for Understanding Long-Form Multimodal Videos
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2025)
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2025)
Tracking with Human-Intent Reasoning
by: Zhu, Jiawen, et al.
Published: (2023)
by: Zhu, Jiawen, et al.
Published: (2023)
Towards Global Localization using Multi-Modal Object-Instance Re-Identification
by: Chavan, Aneesh, et al.
Published: (2024)
by: Chavan, Aneesh, et al.
Published: (2024)
Real-Time Human Reconstruction and Animation using Feed-Forward Gaussian Splatting
by: Chatterjee, Devdoot, et al.
Published: (2026)
by: Chatterjee, Devdoot, et al.
Published: (2026)
AI-Generated Lecture Slides for Improving Slide Element Detection and Retrieval
by: Maniyar, Suyash, et al.
Published: (2025)
by: Maniyar, Suyash, et al.
Published: (2025)
Understanding Task Transfer in Vision-Language Models
by: Sachdeva, Bhuvan, et al.
Published: (2025)
by: Sachdeva, Bhuvan, et al.
Published: (2025)
CCF: Cross Correcting Framework for Pedestrian Trajectory Prediction
by: Chib, Pranav Singh, et al.
Published: (2024)
by: Chib, Pranav Singh, et al.
Published: (2024)
Similar Items
-
Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media
by: M, Megha Mariam K., et al.
Published: (2026) -
Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning
by: Raj, Ankita, et al.
Published: (2025) -
Can Better Text Semantics in Prompt Tuning Improve VLM Generalization?
by: Kuchibhotla, Hari Chandana, et al.
Published: (2024) -
PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction
by: Mishra, Naman, et al.
Published: (2026) -
Source-Free Domain Adaptation by Optimizing Batch-Wise Cosine Similarity
by: Pathak, Harsharaj, et al.
Published: (2026)