Saved in:
| Main Authors: | Khindkar, Vaishnavi, Balasubramanian, Vineeth, Arora, Chetan, Subramanian, Anbumani, Jawahar, C. V. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2411.13302 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media
by: M, Megha Mariam K., et al.
Published: (2026)
by: M, Megha Mariam K., et al.
Published: (2026)
Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning
by: Raj, Ankita, et al.
Published: (2025)
by: Raj, Ankita, et al.
Published: (2025)
Can Better Text Semantics in Prompt Tuning Improve VLM Generalization?
by: Kuchibhotla, Hari Chandana, et al.
Published: (2024)
by: Kuchibhotla, Hari Chandana, et al.
Published: (2024)
PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction
by: Mishra, Naman, et al.
Published: (2026)
by: Mishra, Naman, et al.
Published: (2026)
Source-Free Domain Adaptation by Optimizing Batch-Wise Cosine Similarity
by: Pathak, Harsharaj, et al.
Published: (2026)
by: Pathak, Harsharaj, et al.
Published: (2026)
Pedestrian Intention and Trajectory Prediction in Unstructured Traffic Using IDD-PeD
by: Bokkasam, Ruthvik, et al.
Published: (2025)
by: Bokkasam, Ruthvik, et al.
Published: (2025)
Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
Temporal-contextual Event Learning for Pedestrian Crossing Intent Prediction
by: Liang, Hongbin, et al.
Published: (2025)
by: Liang, Hongbin, et al.
Published: (2025)
Pedestrian Crossing Intent Prediction via Psychological Features and Transformer Fusion
by: Ashayer, Sima, et al.
Published: (2026)
by: Ashayer, Sima, et al.
Published: (2026)
C2FDrone: Coarse-to-Fine Drone-to-Drone Detection using Vision Transformer Networks
by: Rebbapragada, Sairam VC, et al.
Published: (2024)
by: Rebbapragada, Sairam VC, et al.
Published: (2024)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
ACIT: Attention-Guided Cross-Modal Interaction Transformer for Pedestrian Crossing Intention Prediction
by: Li, Yuanzhe, et al.
Published: (2025)
by: Li, Yuanzhe, et al.
Published: (2025)
Towards Global Localization using Multi-Modal Object-Instance Re-Identification
by: Chavan, Aneesh, et al.
Published: (2024)
by: Chavan, Aneesh, et al.
Published: (2024)
ECoDepth: Effective Conditioning of Diffusion Models for Monocular Depth Estimation
by: Patni, Suraj, et al.
Published: (2024)
by: Patni, Suraj, et al.
Published: (2024)
iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
by: Mehrotra, Sarthak, et al.
Published: (2025)
by: Mehrotra, Sarthak, et al.
Published: (2025)
LQ-Adapter: ViT-Adapter with Learnable Queries for Gallbladder Cancer Detection from Ultrasound Image
by: Madan, Chetan, et al.
Published: (2024)
by: Madan, Chetan, et al.
Published: (2024)
$\oslash$ Source Models Leak What They Shouldn't $\nrightarrow$: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimization
by: Devalapally, Arnav, et al.
Published: (2026)
by: Devalapally, Arnav, et al.
Published: (2026)
Precise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class Imbalance
by: Santra, Sanchayan, et al.
Published: (2025)
by: Santra, Sanchayan, et al.
Published: (2025)
Generative Adversarial Patches for Physical Attacks on Cross-Modal Pedestrian Re-Identification
by: Su, Yue, et al.
Published: (2024)
by: Su, Yue, et al.
Published: (2024)
Robust Pedestrian Detection with Uncertain Modality
by: Bie, Qian, et al.
Published: (2026)
by: Bie, Qian, et al.
Published: (2026)
Feature Space Perturbation: A Panacea to Enhanced Transferability Estimation
by: Khoba, Prafful Kumar, et al.
Published: (2025)
by: Khoba, Prafful Kumar, et al.
Published: (2025)
LogicCBMs: Logic-Enhanced Concept-Based Learning
by: Vemuri, Deepika SN, et al.
Published: (2025)
by: Vemuri, Deepika SN, et al.
Published: (2025)
Can Generative Video Models Help Pose Estimation?
by: Cai, Ruojin, et al.
Published: (2024)
by: Cai, Ruojin, et al.
Published: (2024)
IndicSTR12: A Dataset for Indic Scene Text Recognition
by: Lunia, Harsh, et al.
Published: (2024)
by: Lunia, Harsh, et al.
Published: (2024)
Advancing Question Answering on Handwritten Documents: A State-of-the-Art Recognition-Based Model for HW-SQuAD
by: Pal, Aniket, et al.
Published: (2024)
by: Pal, Aniket, et al.
Published: (2024)
Attend to what I say: Highlighting relevant content on slides
by: M, Megha Mariam K, et al.
Published: (2026)
by: M, Megha Mariam K, et al.
Published: (2026)
Source-free Video Domain Adaptation by Learning from Noisy Labels
by: Dasgupta, Avijit, et al.
Published: (2023)
by: Dasgupta, Avijit, et al.
Published: (2023)
Context-aware Multi-task Learning for Pedestrian Intent and Trajectory Prediction
by: Munir, Farzeen, et al.
Published: (2024)
by: Munir, Farzeen, et al.
Published: (2024)
Advancing Ante-Hoc Explainable Models through Generative Adversarial Networks
by: Garg, Tanmay, et al.
Published: (2024)
by: Garg, Tanmay, et al.
Published: (2024)
Open-Set Object Detection By Aligning Known Class Representations
by: Sarkar, Hiran, et al.
Published: (2024)
by: Sarkar, Hiran, et al.
Published: (2024)
On Evaluation of Vision Datasets and Models using Human Competency Frameworks
by: Ramachandran, Rahul, et al.
Published: (2024)
by: Ramachandran, Rahul, et al.
Published: (2024)
Understanding Task Transfer in Vision-Language Models
by: Sachdeva, Bhuvan, et al.
Published: (2025)
by: Sachdeva, Bhuvan, et al.
Published: (2025)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
by: Park, Simon, et al.
Published: (2025)
by: Park, Simon, et al.
Published: (2025)
Prompt2LVideos: Exploring Prompts for Understanding Long-Form Multimodal Videos
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2025)
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2025)
Pedestrian Attribute Recognition via Hierarchical Cross-Modality HyperGraph Learning
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Real-Time Human Reconstruction and Animation using Feed-Forward Gaussian Splatting
by: Chatterjee, Devdoot, et al.
Published: (2026)
by: Chatterjee, Devdoot, et al.
Published: (2026)
AI-Generated Lecture Slides for Improving Slide Element Detection and Retrieval
by: Maniyar, Suyash, et al.
Published: (2025)
by: Maniyar, Suyash, et al.
Published: (2025)
Can Large Pretrained Depth Estimation Models Help With Image Dehazing?
by: Zhang, Hongfei, et al.
Published: (2025)
by: Zhang, Hongfei, et al.
Published: (2025)
Fiducial Focus Augmentation for Facial Landmark Detection
by: Kar, Purbayan, et al.
Published: (2024)
by: Kar, Purbayan, et al.
Published: (2024)
FocusMAE: Gallbladder Cancer Detection from Ultrasound Videos with Focused Masked Autoencoders
by: Basu, Soumen, et al.
Published: (2024)
by: Basu, Soumen, et al.
Published: (2024)
Similar Items
-
Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media
by: M, Megha Mariam K., et al.
Published: (2026) -
Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning
by: Raj, Ankita, et al.
Published: (2025) -
Can Better Text Semantics in Prompt Tuning Improve VLM Generalization?
by: Kuchibhotla, Hari Chandana, et al.
Published: (2024) -
PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction
by: Mishra, Naman, et al.
Published: (2026) -
Source-Free Domain Adaptation by Optimizing Batch-Wise Cosine Similarity
by: Pathak, Harsharaj, et al.
Published: (2026)