Navigating Hallucinations for Reasoning of Unintentional Activities
Fuente:
arXiv
Saved in:
| Main Authors: | Grover, Shresth, Vineet, Vibhav, Rawat, Yogesh S |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
by: Grover, Shresth, et al.
Published: (2025)
by: Grover, Shresth, et al.
Published: (2025)
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
by: Azad, Shehreen, et al.
Published: (2025)
by: Azad, Shehreen, et al.
Published: (2025)
StreamReady: Learning What to Answer and When in Long Streaming Videos
by: Azad, Shehreen, et al.
Published: (2026)
by: Azad, Shehreen, et al.
Published: (2026)
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
by: Kumar, Akash, et al.
Published: (2025)
by: Kumar, Akash, et al.
Published: (2025)
OmViD: Omni-supervised active learning for video action detection
by: Rana, Aayush, et al.
Published: (2025)
by: Rana, Aayush, et al.
Published: (2025)
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
by: Modi, Rajat, et al.
Published: (2024)
by: Modi, Rajat, et al.
Published: (2024)
Understanding Depth and Height Perception in Large Visual-Language Models
by: Azad, Shehreen, et al.
Published: (2024)
by: Azad, Shehreen, et al.
Published: (2024)
Robustness Analysis on Foundational Segmentation Models
by: Schiappa, Madeline Chantry, et al.
Published: (2023)
by: Schiappa, Madeline Chantry, et al.
Published: (2023)
Physics Knowledge in Frontier Models: A Diagnostic Study of Failure Modes
by: Bagdonaviciute, Ieva, et al.
Published: (2025)
by: Bagdonaviciute, Ieva, et al.
Published: (2025)
Activity-Biometrics: Person Identification from Daily Activities
by: Azad, Shehreen, et al.
Published: (2024)
by: Azad, Shehreen, et al.
Published: (2024)
DisenQ: Disentangling Q-Former for Activity-Biometrics
by: Azad, Shehreen, et al.
Published: (2025)
by: Azad, Shehreen, et al.
Published: (2025)
How Do Inpainting Artifacts Propagate to Language?
by: Yashwante, Pratham, et al.
Published: (2026)
by: Yashwante, Pratham, et al.
Published: (2026)
ProDiG: Progressive Diffusion-Guided Gaussian Splatting for Aerial to Ground Reconstruction
by: Mitra, Sirshapan, et al.
Published: (2026)
by: Mitra, Sirshapan, et al.
Published: (2026)
Colors See Colors Ignore: Clothes Changing ReID with Color Disentanglement
by: Pathak, Priyank, et al.
Published: (2025)
by: Pathak, Priyank, et al.
Published: (2025)
DIFFER: Disentangling Identity Features via Semantic Cues for Clothes-Changing Person Re-ID
by: Liang, Xin, et al.
Published: (2025)
by: Liang, Xin, et al.
Published: (2025)
Coarse Attribute Prediction with Task Agnostic Distillation for Real World Clothes Changing ReID
by: Pathak, Priyank, et al.
Published: (2025)
by: Pathak, Priyank, et al.
Published: (2025)
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
by: Ravi, Sahithya, et al.
Published: (2025)
by: Ravi, Sahithya, et al.
Published: (2025)
GaitCrafter: Diffusion Model for Biometric Preserving Gait Synthesis
by: Mitra, Sirshapan, et al.
Published: (2025)
by: Mitra, Sirshapan, et al.
Published: (2025)
STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding
by: Garg, Aaryan, et al.
Published: (2025)
by: Garg, Aaryan, et al.
Published: (2025)
Scaling Open-Vocabulary Action Detection
by: Sia, Zhen Hao, et al.
Published: (2025)
by: Sia, Zhen Hao, et al.
Published: (2025)
PEEKABOO: Interactive Video Generation via Masked-Diffusion
by: Jain, Yash, et al.
Published: (2023)
by: Jain, Yash, et al.
Published: (2023)
Stable Mean Teacher for Semi-supervised Video Action Detection
by: Kumar, Akash, et al.
Published: (2024)
by: Kumar, Akash, et al.
Published: (2024)
MolVision: Molecular Property Prediction with Vision Language Models
by: Adak, Deepan, et al.
Published: (2025)
by: Adak, Deepan, et al.
Published: (2025)
iSafetyBench: A video-language benchmark for safety in industrial environment
by: Abdullah, Raiyaan, et al.
Published: (2025)
by: Abdullah, Raiyaan, et al.
Published: (2025)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
by: Kumar, Akash, et al.
Published: (2025)
by: Kumar, Akash, et al.
Published: (2025)
Asynchronous Perception Machine For Efficient Test-Time-Training
by: Modi, Rajat, et al.
Published: (2024)
by: Modi, Rajat, et al.
Published: (2024)
Grounding Task Assistance with Multimodal Cues from a Single Demonstration
by: Sarch, Gabriel, et al.
Published: (2025)
by: Sarch, Gabriel, et al.
Published: (2025)
Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models
by: Wang, Jiayu, et al.
Published: (2024)
by: Wang, Jiayu, et al.
Published: (2024)
LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models
by: Pathak, Priyank, et al.
Published: (2025)
by: Pathak, Priyank, et al.
Published: (2025)
EZ-CLIP: Efficient Zeroshot Video Action Recognition
by: Ahmad, Shahzad, et al.
Published: (2023)
by: Ahmad, Shahzad, et al.
Published: (2023)
Revealing Unintentional Information Leakage in Low-Dimensional Facial Portrait Representations
by: Anderson, Kathleen, et al.
Published: (2025)
by: Anderson, Kathleen, et al.
Published: (2025)
OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive Tasks
by: Wu, Jing, et al.
Published: (2026)
by: Wu, Jing, et al.
Published: (2026)
MolSight: Molecular Property Prediction with Images
by: Baranwal, Aaditya, et al.
Published: (2026)
by: Baranwal, Aaditya, et al.
Published: (2026)
Advancing Automatic Photovoltaic Defect Detection using Semi-Supervised Semantic Segmentation of Electroluminescence Images
by: Jha, Abhishek, et al.
Published: (2024)
by: Jha, Abhishek, et al.
Published: (2024)
RobustGait: Robustness Analysis for Appearance Based Gait Recognition
by: Sayera, Reeshoon, et al.
Published: (2025)
by: Sayera, Reeshoon, et al.
Published: (2025)
Sky2Ground: A Benchmark for Site Modeling under Varying Altitude
by: Wang, Zengyan, et al.
Published: (2026)
by: Wang, Zengyan, et al.
Published: (2026)
Scaling Vision-and-Language Navigation With Offline RL
by: Bundele, Valay, et al.
Published: (2024)
by: Bundele, Valay, et al.
Published: (2024)
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
Foundation Models for Video Understanding: A Survey
by: Madan, Neelu, et al.
Published: (2024)
by: Madan, Neelu, et al.
Published: (2024)
Similar Items
-
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
by: Grover, Shresth, et al.
Published: (2025) -
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
by: Azad, Shehreen, et al.
Published: (2025) -
StreamReady: Learning What to Answer and When in Long Streaming Videos
by: Azad, Shehreen, et al.
Published: (2026) -
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
by: Kumar, Akash, et al.
Published: (2025) -
OmViD: Omni-supervised active learning for video action detection
by: Rana, Aayush, et al.
Published: (2025)