HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
Fuente:
arXiv
Salvato in:
| Autori principali: | Saito, Kuniaki, Shinoda, Risa, Tanaka, Shohei, Hirasawa, Tosho, Okura, Fumio, Ushiku, Yoshitaka |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
di: Saito, Kuniaki, et al.
Pubblicazione: (2026)
di: Saito, Kuniaki, et al.
Pubblicazione: (2026)
SBS Figures: Pre-training Figure QA from Stage-by-Stage Synthesized Images
di: Shinoda, Risa, et al.
Pubblicazione: (2024)
di: Shinoda, Risa, et al.
Pubblicazione: (2024)
BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment
di: Shinoda, Risa, et al.
Pubblicazione: (2026)
di: Shinoda, Risa, et al.
Pubblicazione: (2026)
CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
di: Saito, Kuniaki, et al.
Pubblicazione: (2025)
di: Saito, Kuniaki, et al.
Pubblicazione: (2025)
AgroBench: Vision-Language Model Benchmark in Agriculture
di: Shinoda, Risa, et al.
Pubblicazione: (2025)
di: Shinoda, Risa, et al.
Pubblicazione: (2025)
SciPostGen: Bridging the Gap between Scientific Papers and Poster Layouts
di: Inadumi, Shun, et al.
Pubblicazione: (2025)
di: Inadumi, Shun, et al.
Pubblicazione: (2025)
GaussianPlant: Structure-aligned Gaussian Splatting for 3D Reconstruction of Plants
di: Yang, Yang, et al.
Pubblicazione: (2025)
di: Yang, Yang, et al.
Pubblicazione: (2025)
COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language Benchmark
di: Maeda, Koki, et al.
Pubblicazione: (2024)
di: Maeda, Koki, et al.
Pubblicazione: (2024)
SciPostLayoutTree: A Dataset for Structural Analysis of Scientific Posters
di: Tanaka, Shohei, et al.
Pubblicazione: (2025)
di: Tanaka, Shohei, et al.
Pubblicazione: (2025)
SciPostLayout: A Dataset for Layout Analysis and Layout Generation of Scientific Posters
di: Tanaka, Shohei, et al.
Pubblicazione: (2024)
di: Tanaka, Shohei, et al.
Pubblicazione: (2024)
PetFace: A Large-Scale Dataset and Benchmark for Animal Identification
di: Shinoda, Risa, et al.
Pubblicazione: (2024)
di: Shinoda, Risa, et al.
Pubblicazione: (2024)
Interaction-via-Actions: Cattle Interaction Detection with Joint Learning of Action-Interaction Latent Space
di: Nakagawa, Ren, et al.
Pubblicazione: (2025)
di: Nakagawa, Ren, et al.
Pubblicazione: (2025)
VidHal: Benchmarking Temporal Hallucinations in Vision LLMs
di: Choong, Wey Yeh, et al.
Pubblicazione: (2024)
di: Choong, Wey Yeh, et al.
Pubblicazione: (2024)
Evaluating the Capability of Video Question Generation for Expert Knowledge Elicitation
di: Zhang, Huaying, et al.
Pubblicazione: (2025)
di: Zhang, Huaying, et al.
Pubblicazione: (2025)
Zero-shot Hierarchical Plant Segmentation via Foundation Segmentation Models and Text-to-image Attention
di: Xing, Junhao, et al.
Pubblicazione: (2025)
di: Xing, Junhao, et al.
Pubblicazione: (2025)
HalCECE: A Framework for Explainable Hallucination Detection through Conceptual Counterfactuals in Image Captioning
di: Lymperaiou, Maria, et al.
Pubblicazione: (2025)
di: Lymperaiou, Maria, et al.
Pubblicazione: (2025)
OpenAnimalTracks: A Dataset for Animal Track Recognition
di: Shinoda, Risa, et al.
Pubblicazione: (2024)
di: Shinoda, Risa, et al.
Pubblicazione: (2024)
CAMOT: Camera Angle-aware Multi-Object Tracking
di: Limanta, Felix, et al.
Pubblicazione: (2024)
di: Limanta, Felix, et al.
Pubblicazione: (2024)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
di: Wang, Xinran, et al.
Pubblicazione: (2026)
di: Wang, Xinran, et al.
Pubblicazione: (2026)
EnsemHalDet: Robust VLM Hallucination Detection via Ensemble of Internal State Detectors
di: Miyazato, Ryuhei, et al.
Pubblicazione: (2026)
di: Miyazato, Ryuhei, et al.
Pubblicazione: (2026)
Gaussian Mesh Renderer for Lightweight Differentiable Rendering
di: Liu, Xinpeng, et al.
Pubblicazione: (2026)
di: Liu, Xinpeng, et al.
Pubblicazione: (2026)
HalLoc: Token-level Localization of Hallucinations for Vision Language Models
di: Park, Eunkyu, et al.
Pubblicazione: (2025)
di: Park, Eunkyu, et al.
Pubblicazione: (2025)
WarrantScore: Modeling Warrants between Claims and Evidence for Substantiation Evaluation in Peer Reviews
di: Mori, Kiyotada, et al.
Pubblicazione: (2026)
di: Mori, Kiyotada, et al.
Pubblicazione: (2026)
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
di: Ohkawa, Takehiko, et al.
Pubblicazione: (2023)
di: Ohkawa, Takehiko, et al.
Pubblicazione: (2023)
LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models
di: Qiu, Han, et al.
Pubblicazione: (2024)
di: Qiu, Han, et al.
Pubblicazione: (2024)
Learning Contrastive Multimodal Fusion with Improved Modality Dropout for Disease Detection and Prediction
di: Gu, Yi, et al.
Pubblicazione: (2025)
di: Gu, Yi, et al.
Pubblicazione: (2025)
Towards Safer Mobile Agents: Scalable Generation and Evaluation of Diverse Scenarios for VLMs
di: Taniguchi, Takara, et al.
Pubblicazione: (2026)
di: Taniguchi, Takara, et al.
Pubblicazione: (2026)
MultiModal Fine-tuning with Synthetic Captions
di: Enomoto, Shohei, et al.
Pubblicazione: (2026)
di: Enomoto, Shohei, et al.
Pubblicazione: (2026)
Unsupervised 3D Human Pose Estimation via Conditional Multi-view Ancestral Sampling
di: Goto, Ryohei, et al.
Pubblicazione: (2026)
di: Goto, Ryohei, et al.
Pubblicazione: (2026)
TreeFormer: Single-view Plant Skeleton Estimation via Tree-constrained Graph Generation
di: Liu, Xinpeng, et al.
Pubblicazione: (2024)
di: Liu, Xinpeng, et al.
Pubblicazione: (2024)
PlantPose: Universal Plant Skeleton Estimation via Tree-constrained Graph Generation
di: Liu, Xinpeng, et al.
Pubblicazione: (2026)
di: Liu, Xinpeng, et al.
Pubblicazione: (2026)
Mitigating Image Captioning Hallucinations in Vision-Language Models
di: Zhao, Fei, et al.
Pubblicazione: (2025)
di: Zhao, Fei, et al.
Pubblicazione: (2025)
Benchmarking and Improving Detail Image Caption
di: Dong, Hongyuan, et al.
Pubblicazione: (2024)
di: Dong, Hongyuan, et al.
Pubblicazione: (2024)
AnimalClue: Recognizing Animals by their Traces
di: Shinoda, Risa, et al.
Pubblicazione: (2025)
di: Shinoda, Risa, et al.
Pubblicazione: (2025)
DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning
di: Wei, Yuancheng, et al.
Pubblicazione: (2026)
di: Wei, Yuancheng, et al.
Pubblicazione: (2026)
HoGS: Unified Near and Far Object Reconstruction via Homogeneous Gaussian Splatting
di: Liu, Xinpeng, et al.
Pubblicazione: (2025)
di: Liu, Xinpeng, et al.
Pubblicazione: (2025)
Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge
di: Xu, Yuzheng, et al.
Pubblicazione: (2026)
di: Xu, Yuzheng, et al.
Pubblicazione: (2026)
DP-SfM: Dual-Pixel Structure-from-Motion without Scale Ambiguity
di: Makabe, Lilika, et al.
Pubblicazione: (2026)
di: Makabe, Lilika, et al.
Pubblicazione: (2026)
Spectral Sensitivity Estimation with an Uncalibrated Diffraction Grating
di: Makabe, Lilika, et al.
Pubblicazione: (2025)
di: Makabe, Lilika, et al.
Pubblicazione: (2025)
Near-light Photometric Stereo with Symmetric Lights
di: Makabe, Lilika, et al.
Pubblicazione: (2026)
di: Makabe, Lilika, et al.
Pubblicazione: (2026)
Documenti analoghi
-
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
di: Saito, Kuniaki, et al.
Pubblicazione: (2026) -
SBS Figures: Pre-training Figure QA from Stage-by-Stage Synthesized Images
di: Shinoda, Risa, et al.
Pubblicazione: (2024) -
BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment
di: Shinoda, Risa, et al.
Pubblicazione: (2026) -
CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
di: Saito, Kuniaki, et al.
Pubblicazione: (2025) -
AgroBench: Vision-Language Model Benchmark in Agriculture
di: Shinoda, Risa, et al.
Pubblicazione: (2025)