DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xinran, Zhang, Yuxuan, Zhang, Xiao, Yan, Haolong, Diao, Muxi, Xu, Songyu, Yan, Zhonghao, Li, Hongbing, Liang, Kongming, Ma, Zhanyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation
von: Wang, Xinran, et al.
Veröffentlicht: (2025)
von: Wang, Xinran, et al.
Veröffentlicht: (2025)
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
von: Wang, Xinran, et al.
Veröffentlicht: (2025)
von: Wang, Xinran, et al.
Veröffentlicht: (2025)
From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
von: Wang, Xinran, et al.
Veröffentlicht: (2024)
von: Wang, Xinran, et al.
Veröffentlicht: (2024)
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
Mitigating Image Captioning Hallucinations in Vision-Language Models
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
Hepato-LLaVA: An Expert MLLM with Sparse Topo-Pack Attention for Hepatocellular Pathology Analysis on Whole Slide Images
von: Yang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Yang, Yuxuan, et al.
Veröffentlicht: (2026)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
Hallucination Localization in Video Captioning
von: Nakada, Shota, et al.
Veröffentlicht: (2025)
von: Nakada, Shota, et al.
Veröffentlicht: (2025)
RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs
von: Xu, Pengju, et al.
Veröffentlicht: (2025)
von: Xu, Pengju, et al.
Veröffentlicht: (2025)
Multi-Scale and Detail-Enhanced Segment Anything Model for Salient Object Detection
von: Gao, Shixuan, et al.
Veröffentlicht: (2024)
von: Gao, Shixuan, et al.
Veröffentlicht: (2024)
MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision
von: Yan, Zhonghao, et al.
Veröffentlicht: (2025)
von: Yan, Zhonghao, et al.
Veröffentlicht: (2025)
Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
von: Wu, Haoning, et al.
Veröffentlicht: (2023)
von: Wu, Haoning, et al.
Veröffentlicht: (2023)
Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting
von: Diao, Muxi, et al.
Veröffentlicht: (2026)
von: Diao, Muxi, et al.
Veröffentlicht: (2026)
HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
von: Zhang, Zijian, et al.
Veröffentlicht: (2025)
von: Zhang, Zijian, et al.
Veröffentlicht: (2025)
Dental-TriageBench: Benchmarking Multimodal Reasoning for Hierarchical Dental Triage
von: He, Ziyi, et al.
Veröffentlicht: (2026)
von: He, Ziyi, et al.
Veröffentlicht: (2026)
VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics
von: Yin, Zhiyu, et al.
Veröffentlicht: (2026)
von: Yin, Zhiyu, et al.
Veröffentlicht: (2026)
PolySmart @ TRECVid 2024 Video Captioning (VTT)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
Towards Universal Modal Tracking with Online Dense Temporal Token Learning
von: Zheng, Yaozong, et al.
Veröffentlicht: (2025)
von: Zheng, Yaozong, et al.
Veröffentlicht: (2025)
RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos
von: Yang, Zixi, et al.
Veröffentlicht: (2025)
von: Yang, Zixi, et al.
Veröffentlicht: (2025)
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
von: Ge, Shiping, et al.
Veröffentlicht: (2024)
von: Ge, Shiping, et al.
Veröffentlicht: (2024)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
Video-Mediated Emotion Disclosure: Expressions of Fear, Sadness, and Joy by People with Schizophrenia on YouTube
von: Liu, Jiaying Lizzy, et al.
Veröffentlicht: (2025)
von: Liu, Jiaying Lizzy, et al.
Veröffentlicht: (2025)
PetalView: Fine-grained Location and Orientation Extraction of Street-view Images via Cross-view Local Search with Supplementary Materials
von: Hu, Wenmiao, et al.
Veröffentlicht: (2024)
von: Hu, Wenmiao, et al.
Veröffentlicht: (2024)
Differentiable JPEG: The Devil is in the Details
von: Reich, Christoph, et al.
Veröffentlicht: (2023)
von: Reich, Christoph, et al.
Veröffentlicht: (2023)
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
von: Wang, Qiang, et al.
Veröffentlicht: (2025)
von: Wang, Qiang, et al.
Veröffentlicht: (2025)
PaveBench: A Versatile Benchmark for Pavement Distress Perception and Interactive Vision-Language Analysis
von: Li, Dexiang, et al.
Veröffentlicht: (2026)
von: Li, Dexiang, et al.
Veröffentlicht: (2026)
LoginMEA: Local-to-Global Interaction Network for Multi-modal Entity Alignment
von: Su, Taoyu, et al.
Veröffentlicht: (2024)
von: Su, Taoyu, et al.
Veröffentlicht: (2024)
Benchmarking and Improving LVLMs on Event Extraction from Multimedia Documents
von: Xing, Fuyu, et al.
Veröffentlicht: (2025)
von: Xing, Fuyu, et al.
Veröffentlicht: (2025)
Edit As You Wish: Video Caption Editing with Multi-grained User Control
von: Yao, Linli, et al.
Veröffentlicht: (2023)
von: Yao, Linli, et al.
Veröffentlicht: (2023)
Less for More: Enhanced Feedback-aligned Mixed LLMs for Molecule Caption Generation and Fine-Grained NLI Evaluation
von: Gkoumas, Dimitris, et al.
Veröffentlicht: (2024)
von: Gkoumas, Dimitris, et al.
Veröffentlicht: (2024)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
von: Yan, Xin, et al.
Veröffentlicht: (2024)
von: Yan, Xin, et al.
Veröffentlicht: (2024)
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
von: Li, Huilai, et al.
Veröffentlicht: (2025)
von: Li, Huilai, et al.
Veröffentlicht: (2025)
See or Guess: Counterfactually Regularized Image Captioning
von: Cao, Qian, et al.
Veröffentlicht: (2024)
von: Cao, Qian, et al.
Veröffentlicht: (2024)
DeepfakeBench-MM: A Comprehensive Benchmark for Multimodal Deepfake Detection
von: Zhao, Kangran, et al.
Veröffentlicht: (2025)
von: Zhao, Kangran, et al.
Veröffentlicht: (2025)
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
Human-Machine Collaboration-Guided Space Design: Combination of Machine Learning Models and Humanistic Design Concepts
von: Yang, Yuxuan
Veröffentlicht: (2025)
von: Yang, Yuxuan
Veröffentlicht: (2025)
SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation
von: Wang, Xinran, et al.
Veröffentlicht: (2025) -
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
von: Wang, Xinran, et al.
Veröffentlicht: (2025) -
From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
von: Wang, Xinran, et al.
Veröffentlicht: (2024) -
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
von: Ma, Ziyang, et al.
Veröffentlicht: (2025) -
Mitigating Image Captioning Hallucinations in Vision-Language Models
von: Zhao, Fei, et al.
Veröffentlicht: (2025)