HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gautam, Sushant, Riegler, Michael A., Halvorsen, Pål |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
von: Gautam, Sushant, et al.
Veröffentlicht: (2026)
von: Gautam, Sushant, et al.
Veröffentlicht: (2026)
DSER: Spectral Epipolar Representation for Efficient Light Field Depth Estimation
von: Mohammad, Noor Islam S., et al.
Veröffentlicht: (2025)
von: Mohammad, Noor Islam S., et al.
Veröffentlicht: (2025)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
von: Zinnen, Mathias, et al.
Veröffentlicht: (2025)
von: Zinnen, Mathias, et al.
Veröffentlicht: (2025)
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
von: Chen, Jingkun, et al.
Veröffentlicht: (2025)
von: Chen, Jingkun, et al.
Veröffentlicht: (2025)
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
von: Su, Yuetong, et al.
Veröffentlicht: (2025)
von: Su, Yuetong, et al.
Veröffentlicht: (2025)
Hierarchical Spatial Algorithms for High-Resolution Image Quantization and Feature Extraction
von: Mohammad, Noor Islam S.
Veröffentlicht: (2025)
von: Mohammad, Noor Islam S.
Veröffentlicht: (2025)
VDPP: Video Depth Post-Processing for Speed and Scalability
von: Yoon, Daewon, et al.
Veröffentlicht: (2026)
von: Yoon, Daewon, et al.
Veröffentlicht: (2026)
OpenFusion++: An Open-vocabulary Real-time Scene Understanding System
von: Jin, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Jin, Xiaofeng, et al.
Veröffentlicht: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
Corn Ear Detection and Orientation Estimation Using Deep Learning
von: Sprague, Nathan, et al.
Veröffentlicht: (2024)
von: Sprague, Nathan, et al.
Veröffentlicht: (2024)
Hierarchical Point-Patch Fusion with Adaptive Patch Codebook for 3D Shape Anomaly Detection
von: Kang, Xueyang, et al.
Veröffentlicht: (2026)
von: Kang, Xueyang, et al.
Veröffentlicht: (2026)
Gr-IoU: Ground-Intersection over Union for Robust Multi-Object Tracking with 3D Geometric Constraints
von: Toida, Keisuke, et al.
Veröffentlicht: (2024)
von: Toida, Keisuke, et al.
Veröffentlicht: (2024)
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
von: Zhang, Junbin, et al.
Veröffentlicht: (2022)
von: Zhang, Junbin, et al.
Veröffentlicht: (2022)
UVLM: A Universal Vision-Language Model Loader for Reproducible Multimodal Benchmarking
von: Perez, Joan, et al.
Veröffentlicht: (2026)
von: Perez, Joan, et al.
Veröffentlicht: (2026)
A large-scale, physically-based synthetic dataset for satellite pose estimation
von: Velkei, Szabolcs, et al.
Veröffentlicht: (2025)
von: Velkei, Szabolcs, et al.
Veröffentlicht: (2025)
ShapBPT: Image Feature Attributions Using Data-Aware Binary Partition Trees
von: Rashid, Muhammad, et al.
Veröffentlicht: (2026)
von: Rashid, Muhammad, et al.
Veröffentlicht: (2026)
Medico 2025: Visual Question Answering for Gastrointestinal Imaging
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
Do Generative Metrics Predict YOLO Performance? An Evaluation Across Models, Augmentation Ratios, and Dataset Complexity
von: Marian, Vasile, et al.
Veröffentlicht: (2026)
von: Marian, Vasile, et al.
Veröffentlicht: (2026)
Sat-JEPA-Diff: Bridging Self-Supervised Learning and Generative Diffusion for Remote Sensing
von: Komurcu, Kursat, et al.
Veröffentlicht: (2026)
von: Komurcu, Kursat, et al.
Veröffentlicht: (2026)
3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model
von: Ko, Hyun-kyu, et al.
Veröffentlicht: (2026)
von: Ko, Hyun-kyu, et al.
Veröffentlicht: (2026)
VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors
von: Lyu, Wenbo, et al.
Veröffentlicht: (2025)
von: Lyu, Wenbo, et al.
Veröffentlicht: (2025)
Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling
von: Jung, Seoik, et al.
Veröffentlicht: (2025)
von: Jung, Seoik, et al.
Veröffentlicht: (2025)
AQFusionNet: Multimodal Deep Learning for Air Quality Index Prediction with Imagery and Sensor Data
von: Kushal, Koushik Ahmed, et al.
Veröffentlicht: (2025)
von: Kushal, Koushik Ahmed, et al.
Veröffentlicht: (2025)
Real Time Human Detection by Unmanned Aerial Vehicles
von: Guettala, Walid, et al.
Veröffentlicht: (2024)
von: Guettala, Walid, et al.
Veröffentlicht: (2024)
WildfireVLM: AI-powered Analysis for Early Wildfire Detection and Risk Assessment Using Satellite Imagery
von: Ayanzadeh, Aydin, et al.
Veröffentlicht: (2026)
von: Ayanzadeh, Aydin, et al.
Veröffentlicht: (2026)
Experimental Evaluation of Road-Crossing Decisions by Autonomous Wheelchairs against Environmental Factors
von: Corradini, Franca, et al.
Veröffentlicht: (2024)
von: Corradini, Franca, et al.
Veröffentlicht: (2024)
Polarization-Based Eye Tracking with Personalized Siamese Architectures
von: Kalkanli, Beyza, et al.
Veröffentlicht: (2026)
von: Kalkanli, Beyza, et al.
Veröffentlicht: (2026)
BG-YOLO: A Bidirectional-Guided Method for Underwater Object Detection
von: Zhang, Jian, et al.
Veröffentlicht: (2024)
von: Zhang, Jian, et al.
Veröffentlicht: (2024)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
von: Qesaraku, Bjorna, et al.
Veröffentlicht: (2025)
von: Qesaraku, Bjorna, et al.
Veröffentlicht: (2025)
Image-based Facial Rig Inversion
von: Yang, Tianxiang, et al.
Veröffentlicht: (2025)
von: Yang, Tianxiang, et al.
Veröffentlicht: (2025)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
von: Yanambakkam, Hemanth Teja, et al.
Veröffentlicht: (2025)
von: Yanambakkam, Hemanth Teja, et al.
Veröffentlicht: (2025)
ARTPS: Depth-Enhanced Hybrid Anomaly Detection and Learnable Curiosity Score for Autonomous Rover Target Prioritization
von: Baydemir, Poyraz
Veröffentlicht: (2025)
von: Baydemir, Poyraz
Veröffentlicht: (2025)
TRACES: Temporal Recall with Contextual Embeddings for Real-Time Video Anomaly Detection
von: Siddiqui, Yousuf Ahmed, et al.
Veröffentlicht: (2025)
von: Siddiqui, Yousuf Ahmed, et al.
Veröffentlicht: (2025)
Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
Splat and Distill: Augmenting Teachers with Feed-Forward 3D Reconstruction For 3D-Aware Distillation
von: Shavin, David, et al.
Veröffentlicht: (2026)
von: Shavin, David, et al.
Veröffentlicht: (2026)
Dual-sensing driving detection model
von: K, Leon C. C., et al.
Veröffentlicht: (2025)
von: K, Leon C. C., et al.
Veröffentlicht: (2025)
μ-Net: A Deep Learning-Based Architecture for μ-CT Segmentation
von: Bruno, Pierangela, et al.
Veröffentlicht: (2024)
von: Bruno, Pierangela, et al.
Veröffentlicht: (2024)
Can Local Vision-Language Models improve Activity Recognition over Vision Transformers? -- Case Study on Newborn Resuscitation
von: Guerriero, Enrico, et al.
Veröffentlicht: (2026)
von: Guerriero, Enrico, et al.
Veröffentlicht: (2026)
YOLO Ensemble for UAV-based Multispectral Defect Detection in Wind Turbine Components
von: Svystun, Serhii, et al.
Veröffentlicht: (2025)
von: Svystun, Serhii, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
von: Gautam, Sushant, et al.
Veröffentlicht: (2025) -
VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
von: Gautam, Sushant, et al.
Veröffentlicht: (2026) -
DSER: Spectral Epipolar Representation for Efficient Light Field Depth Estimation
von: Mohammad, Noor Islam S., et al.
Veröffentlicht: (2025) -
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
von: Zinnen, Mathias, et al.
Veröffentlicht: (2025) -
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
von: Chen, Jingkun, et al.
Veröffentlicht: (2025)