Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Vellamcheti, Shanmukha, Kundu, Sanjoy, Aakur, Sathyanarayanan N. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Probabilistic Jump-Diffusion Framework for Open-World Egocentric Activity Recognition
by: Kundu, Sanjoy, et al.
Published: (2025)
by: Kundu, Sanjoy, et al.
Published: (2025)
ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity Recognition
by: Kundu, Sanjoy, et al.
Published: (2025)
by: Kundu, Sanjoy, et al.
Published: (2025)
ALGO: Object-Grounded Visual Commonsense Reasoning for Open-World Egocentric Action Recognition
by: Kundu, Sanjoy, et al.
Published: (2024)
by: Kundu, Sanjoy, et al.
Published: (2024)
Discovering Novel Actions from Open World Egocentric Videos with Object-Grounded Visual Commonsense Reasoning
by: Kundu, Sanjoy, et al.
Published: (2023)
by: Kundu, Sanjoy, et al.
Published: (2023)
CVT-Bench: Counterfactual Viewpoint Transformations Reveal Unstable Spatial Representations in Multimodal LLMs
by: Vellamcheti, Shanmukha, et al.
Published: (2026)
by: Vellamcheti, Shanmukha, et al.
Published: (2026)
CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
EASE: Embodied Active Event Perception via Self-Supervised Energy Minimization
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
Self-supervised Multi-actor Social Activity Understanding in Streaming Videos
by: Trehan, Shubham, et al.
Published: (2024)
by: Trehan, Shubham, et al.
Published: (2024)
Generalized Event Partonomy Inference with Structured Hierarchical Predictive Learning
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
STaTS: Structure-Aware Temporal Sequence Summarization via Statistical Window Merging
by: Bhowmick, Disharee, et al.
Published: (2025)
by: Bhowmick, Disharee, et al.
Published: (2025)
FSP-DETR: Few-Shot Prototypical Parasitic Ova Detection
by: Trehan, Shubham, et al.
Published: (2025)
by: Trehan, Shubham, et al.
Published: (2025)
Capturing Temporal Components for Time Series Classification
by: Vavilthota, Venkata Ragavendra, et al.
Published: (2024)
by: Vavilthota, Venkata Ragavendra, et al.
Published: (2024)
Seeing to Ground: Visual Attention for Hallucination-Resilient MDLLMs
by: Narnaware, Vishal, et al.
Published: (2026)
by: Narnaware, Vishal, et al.
Published: (2026)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
by: Guo, Pinxue, et al.
Published: (2025)
by: Guo, Pinxue, et al.
Published: (2025)
Open-Text Aerial Detection: A Unified Framework For Aerial Visual Grounding And Detection
by: Wei, Guoting, et al.
Published: (2026)
by: Wei, Guoting, et al.
Published: (2026)
METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection
by: Wang, Yongqi, et al.
Published: (2025)
by: Wang, Yongqi, et al.
Published: (2025)
VIHD: Visual Intervention-based Hallucination Detection for Medical Visual Question Answering
by: Chen, Jiayi, et al.
Published: (2026)
by: Chen, Jiayi, et al.
Published: (2026)
Multi-Modal Hallucination Control by Visual Information Grounding
by: Favero, Alessandro, et al.
Published: (2024)
by: Favero, Alessandro, et al.
Published: (2024)
Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
by: Ghosh, Sreyan, et al.
Published: (2024)
by: Ghosh, Sreyan, et al.
Published: (2024)
EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models
by: Villa, Andrés, et al.
Published: (2025)
by: Villa, Andrés, et al.
Published: (2025)
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
by: Jiang, Yubo, et al.
Published: (2026)
by: Jiang, Yubo, et al.
Published: (2026)
V-Loop: Visual Logical Loop Verification for Hallucination Detection in Medical Visual Question Answering
by: Jin, Mengyuan, et al.
Published: (2026)
by: Jin, Mengyuan, et al.
Published: (2026)
GroundCount: Grounding Vision-Language Models with Object Detection for Mitigating Counting Hallucinations
by: Chen, Boyuan, et al.
Published: (2026)
by: Chen, Boyuan, et al.
Published: (2026)
NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving
by: Li, Fuhao, et al.
Published: (2025)
by: Li, Fuhao, et al.
Published: (2025)
CalibFree: Self-Supervised View Feature Separation for Calibration-Free Multi-Camera Multi-Object Tracking
by: Xian, Ruiqi, et al.
Published: (2026)
by: Xian, Ruiqi, et al.
Published: (2026)
AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
by: Galougah, Siminfar Samakoush, et al.
Published: (2025)
by: Galougah, Siminfar Samakoush, et al.
Published: (2025)
SAGE: Sink-Aware Grounded Decoding for Multimodal Hallucination Mitigation
by: Shukla, Tripti, et al.
Published: (2026)
by: Shukla, Tripti, et al.
Published: (2026)
TruthLens: Visual Grounding for Universal DeepFake Reasoning
by: Kundu, Rohit, et al.
Published: (2025)
by: Kundu, Rohit, et al.
Published: (2025)
Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language Models
by: Wu, Mingrui, et al.
Published: (2024)
by: Wu, Mingrui, et al.
Published: (2024)
Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
by: Luo, Meng, et al.
Published: (2025)
by: Luo, Meng, et al.
Published: (2025)
VGS-Decoding: Visual Grounding Score Guided Decoding for Hallucination Mitigation in Medical VLMs
by: Kolli, Govinda, et al.
Published: (2026)
by: Kolli, Govinda, et al.
Published: (2026)
VL-UniTrack: A Unified Framework with Visual-Language Prompts for UAV-Ground Visual Tracking
by: Xu, Boyue, et al.
Published: (2026)
by: Xu, Boyue, et al.
Published: (2026)
Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning
by: Wei, Zhaoyang, et al.
Published: (2025)
by: Wei, Zhaoyang, et al.
Published: (2025)
Pensieve: Retrospect-then-Compare Mitigates Visual Hallucination
by: Yang, Dingchen, et al.
Published: (2024)
by: Yang, Dingchen, et al.
Published: (2024)
A Survey of Multimodal Hallucination Evaluation and Detection
by: Chen, Zhiyuan, et al.
Published: (2025)
by: Chen, Zhiyuan, et al.
Published: (2025)
UniDGF: A Unified Detection-to-Generation Framework for Hierarchical Object Visual Recognition
by: Nan, Xinyu, et al.
Published: (2025)
by: Nan, Xinyu, et al.
Published: (2025)
Localizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs
by: Nguyen, Dung, et al.
Published: (2025)
by: Nguyen, Dung, et al.
Published: (2025)
Improving Generalized Visual Grounding with Instance-aware Joint Learning
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
Learning Visual Grounding from Generative Vision and Language Model
by: Wang, Shijie, et al.
Published: (2024)
by: Wang, Shijie, et al.
Published: (2024)
Hallucination Early Detection in Diffusion Models
by: Betti, Federico, et al.
Published: (2026)
by: Betti, Federico, et al.
Published: (2026)
Similar Items
-
A Probabilistic Jump-Diffusion Framework for Open-World Egocentric Activity Recognition
by: Kundu, Sanjoy, et al.
Published: (2025) -
ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity Recognition
by: Kundu, Sanjoy, et al.
Published: (2025) -
ALGO: Object-Grounded Visual Commonsense Reasoning for Open-World Egocentric Action Recognition
by: Kundu, Sanjoy, et al.
Published: (2024) -
Discovering Novel Actions from Open World Egocentric Videos with Object-Grounded Visual Commonsense Reasoning
by: Kundu, Sanjoy, et al.
Published: (2023) -
CVT-Bench: Counterfactual Viewpoint Transformations Reveal Unstable Spatial Representations in Multimodal LLMs
by: Vellamcheti, Shanmukha, et al.
Published: (2026)