VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
Fuente:
arXiv
Saved in:
| Main Authors: | Gautam, Sushant, Midoglu, Cise, Thambawita, Vajira, Riegler, Michael A., Halvorsen, Pål |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
Medico 2025: Visual Question Answering for Gastrointestinal Imaging
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
TRACES: Temporal Recall with Contextual Embeddings for Real-Time Video Anomaly Detection
by: Siddiqui, Yousuf Ahmed, et al.
Published: (2025)
by: Siddiqui, Yousuf Ahmed, et al.
Published: (2025)
The Impact of Image Resolution on Face Detection: A Comparative Analysis of MTCNN, YOLOv XI and YOLOv XII models
by: Ömercikoğlu, Ahmet Can, et al.
Published: (2025)
by: Ömercikoğlu, Ahmet Can, et al.
Published: (2025)
CADE 2.5 - ZeResFDG: Frequency-Decoupled, Rescaled and Zero-Projected Guidance for SD/SDXL Latent Diffusion Models
by: Rychkovskiy, Denis
Published: (2025)
by: Rychkovskiy, Denis
Published: (2025)
Beyond RGB: Leveraging Vision Transformers for Thermal Weapon Segmentation
by: Kambhatla, Akhila, et al.
Published: (2025)
by: Kambhatla, Akhila, et al.
Published: (2025)
Rethinking Visual Intelligence: Insights from Video Pretraining
by: Acuaviva, Pablo, et al.
Published: (2025)
by: Acuaviva, Pablo, et al.
Published: (2025)
BG-YOLO: A Bidirectional-Guided Method for Underwater Object Detection
by: Zhang, Jian, et al.
Published: (2024)
by: Zhang, Jian, et al.
Published: (2024)
A deep learning approach to track eye movements based on events
by: Seth, Chirag, et al.
Published: (2025)
by: Seth, Chirag, et al.
Published: (2025)
Transforming faces into video stories -- VideoFace2.0
by: Brkljač, Branko, et al.
Published: (2025)
by: Brkljač, Branko, et al.
Published: (2025)
Fairness Without Labels: Pseudo-Balancing for Bias Mitigation in Face Gender Classification
by: Dong, Haohua, et al.
Published: (2025)
by: Dong, Haohua, et al.
Published: (2025)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
by: Patel, Urjitkumar, et al.
Published: (2025)
by: Patel, Urjitkumar, et al.
Published: (2025)
Prompt to Polyp: Medical Text-Conditioned Image Synthesis with Diffusion Models
by: Chaichuk, Mikhail, et al.
Published: (2025)
by: Chaichuk, Mikhail, et al.
Published: (2025)
Explaining What Machines See: XAI Strategies in Deep Object Detection Models
by: Seyedmomeni, FatemehSadat, et al.
Published: (2025)
by: Seyedmomeni, FatemehSadat, et al.
Published: (2025)
Feature Based Methods in Domain Adaptation for Object Detection: A Review Paper
by: Mohamadi, Helia, et al.
Published: (2024)
by: Mohamadi, Helia, et al.
Published: (2024)
A Novel Compression Framework for YOLOv8: Achieving Real-Time Aerial Object Detection on Edge Devices via Structured Pruning and Channel-Wise Distillation
by: Sabaghian, Melika, et al.
Published: (2025)
by: Sabaghian, Melika, et al.
Published: (2025)
Architectural Insights into Knowledge Distillation for Object Detection: A Comprehensive Review
by: Golizadeh, Mahdi, et al.
Published: (2025)
by: Golizadeh, Mahdi, et al.
Published: (2025)
Enhancing Small Object Detection with YOLO: A Novel Framework for Improved Accuracy and Efficiency
by: Moghadami, Mahila, et al.
Published: (2025)
by: Moghadami, Mahila, et al.
Published: (2025)
Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis
by: Korolkov, Vasilii
Published: (2025)
by: Korolkov, Vasilii
Published: (2025)
QSilk: Micrograin Stabilization and Adaptive Quantile Clipping for Detail-Friendly Latent Diffusion
by: Rychkovskiy, Denis
Published: (2025)
by: Rychkovskiy, Denis
Published: (2025)
Polarization-Based Eye Tracking with Personalized Siamese Architectures
by: Kalkanli, Beyza, et al.
Published: (2026)
by: Kalkanli, Beyza, et al.
Published: (2026)
Sat-JEPA-Diff: Bridging Self-Supervised Learning and Generative Diffusion for Remote Sensing
by: Komurcu, Kursat, et al.
Published: (2026)
by: Komurcu, Kursat, et al.
Published: (2026)
RailSafeNet: Visual Scene Understanding for Tram Safety
by: Valach, Ondřej, et al.
Published: (2025)
by: Valach, Ondřej, et al.
Published: (2025)
μ-Net: A Deep Learning-Based Architecture for μ-CT Segmentation
by: Bruno, Pierangela, et al.
Published: (2024)
by: Bruno, Pierangela, et al.
Published: (2024)
YOLO Ensemble for UAV-based Multispectral Defect Detection in Wind Turbine Components
by: Svystun, Serhii, et al.
Published: (2025)
by: Svystun, Serhii, et al.
Published: (2025)
Archival Faces: Detection of Faces in Digitized Historical Documents
by: Vaško, Marek, et al.
Published: (2025)
by: Vaško, Marek, et al.
Published: (2025)
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
by: Zhang, Junbin, et al.
Published: (2022)
by: Zhang, Junbin, et al.
Published: (2022)
AQFusionNet: Multimodal Deep Learning for Air Quality Index Prediction with Imagery and Sensor Data
by: Kushal, Koushik Ahmed, et al.
Published: (2025)
by: Kushal, Koushik Ahmed, et al.
Published: (2025)
ARTPS: Depth-Enhanced Hybrid Anomaly Detection and Learnable Curiosity Score for Autonomous Rover Target Prioritization
by: Baydemir, Poyraz
Published: (2025)
by: Baydemir, Poyraz
Published: (2025)
ROI-GS: Interest-based Local Quality 3D Gaussian Splatting
by: Bui, Quoc-Anh, et al.
Published: (2025)
by: Bui, Quoc-Anh, et al.
Published: (2025)
ROI-NeRFs: Hi-Fi Visualization of Objects of Interest within a Scene by NeRFs Composition
by: Bui, Quoc-Anh, et al.
Published: (2025)
by: Bui, Quoc-Anh, et al.
Published: (2025)
Dense Video Understanding with Gated Residual Tokenization
by: Zhang, Haichao, et al.
Published: (2025)
by: Zhang, Haichao, et al.
Published: (2025)
Unlocking UML Class Diagram Understanding in Vision Language Models
by: Naboichenko, Artem, et al.
Published: (2026)
by: Naboichenko, Artem, et al.
Published: (2026)
Attention Gathers, MLPs Compose: A Causal Analysis of an Action-Outcome Circuit in VideoViT
by: Chereddy, Sai V R
Published: (2026)
by: Chereddy, Sai V R
Published: (2026)
Image-based Facial Rig Inversion
by: Yang, Tianxiang, et al.
Published: (2025)
by: Yang, Tianxiang, et al.
Published: (2025)
EventFlow: Real-Time Neuromorphic Event-Driven Classification of Two-Phase Boiling Flow Regimes
by: Chang, Sanghyeon, et al.
Published: (2025)
by: Chang, Sanghyeon, et al.
Published: (2025)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
by: Qesaraku, Bjorna, et al.
Published: (2025)
by: Qesaraku, Bjorna, et al.
Published: (2025)
Adversarial Patch Attacks on Vision-Based Cargo Occupancy Estimation via Differentiable 3D Simulation
by: Hedna, Mohamed Rissal, et al.
Published: (2025)
by: Hedna, Mohamed Rissal, et al.
Published: (2025)
Similar Items
-
HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025) -
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025) -
SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding
by: Gautam, Sushant, et al.
Published: (2025) -
Medico 2025: Visual Question Answering for Gastrointestinal Imaging
by: Gautam, Sushant, et al.
Published: (2025) -
TRACES: Temporal Recall with Contextual Embeddings for Real-Time Video Anomaly Detection
by: Siddiqui, Yousuf Ahmed, et al.
Published: (2025)