VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Park, Seongheon, Oh, Changdae, Choi, Hyeong Kyu, Du, Sean, Li, Sharon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
di: Park, Seongheon, et al.
Pubblicazione: (2025)
di: Park, Seongheon, et al.
Pubblicazione: (2025)
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
di: Kundu, Souvik, et al.
Pubblicazione: (2025)
di: Kundu, Souvik, et al.
Pubblicazione: (2025)
Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding
di: Cho, Beomsik, et al.
Pubblicazione: (2025)
di: Cho, Beomsik, et al.
Pubblicazione: (2025)
Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
di: Padhi, Trilok, et al.
Pubblicazione: (2025)
di: Padhi, Trilok, et al.
Pubblicazione: (2025)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
di: Lu, Yujie, et al.
Pubblicazione: (2024)
di: Lu, Yujie, et al.
Pubblicazione: (2024)
Contrastive Language Prompting to Ease False Positives in Medical Anomaly Detection
di: Park, YeongHyeon, et al.
Pubblicazione: (2024)
di: Park, YeongHyeon, et al.
Pubblicazione: (2024)
KorMedMCQA-V: A Multimodal Benchmark for Evaluating Vision-Language Models on the Korean Medical Licensing Examination
di: Choi, Byungjin, et al.
Pubblicazione: (2026)
di: Choi, Byungjin, et al.
Pubblicazione: (2026)
Enhancing Temporal Action Localization: Advanced S6 Modeling with Recurrent Mechanism
di: Lee, Sangyoun, et al.
Pubblicazione: (2024)
di: Lee, Sangyoun, et al.
Pubblicazione: (2024)
MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
di: Deng, Jingyuan, et al.
Pubblicazione: (2025)
di: Deng, Jingyuan, et al.
Pubblicazione: (2025)
ModeX: Evaluator-Free Best-of-N Selection for Open-Ended Generation
di: Choi, Hyeong Kyu, et al.
Pubblicazione: (2026)
di: Choi, Hyeong Kyu, et al.
Pubblicazione: (2026)
Uncertainty-Aware Evaluation for Vision-Language Models
di: Kostumov, Vasily, et al.
Pubblicazione: (2024)
di: Kostumov, Vasily, et al.
Pubblicazione: (2024)
LVLM-Aided Alignment of Task-Specific Vision Models
di: Koebler, Alexander, et al.
Pubblicazione: (2025)
di: Koebler, Alexander, et al.
Pubblicazione: (2025)
TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models
di: Shao, Wenqi, et al.
Pubblicazione: (2023)
di: Shao, Wenqi, et al.
Pubblicazione: (2023)
Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation
di: Mao, Jiawei, et al.
Pubblicazione: (2026)
di: Mao, Jiawei, et al.
Pubblicazione: (2026)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
di: Hwang, Yerin, et al.
Pubblicazione: (2025)
di: Hwang, Yerin, et al.
Pubblicazione: (2025)
ClaudesLens: Uncertainty Quantification in Computer Vision Models
di: Shaar, Mohamad Al, et al.
Pubblicazione: (2024)
di: Shaar, Mohamad Al, et al.
Pubblicazione: (2024)
On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models
di: Seo, Hoigi, et al.
Pubblicazione: (2025)
di: Seo, Hoigi, et al.
Pubblicazione: (2025)
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
di: Chandu, Khyathi Raghavi, et al.
Pubblicazione: (2024)
di: Chandu, Khyathi Raghavi, et al.
Pubblicazione: (2024)
HAWAII: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language Models
di: Wang, Yimu, et al.
Pubblicazione: (2025)
di: Wang, Yimu, et al.
Pubblicazione: (2025)
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models
di: He, Zoe Wanying, et al.
Pubblicazione: (2025)
di: He, Zoe Wanying, et al.
Pubblicazione: (2025)
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
di: Zhao, Xiaohan, et al.
Pubblicazione: (2026)
di: Zhao, Xiaohan, et al.
Pubblicazione: (2026)
LVLM-COUNT: Enhancing the Counting Ability of Large Vision-Language Models
di: Qharabagh, Muhammad Fetrat, et al.
Pubblicazione: (2024)
di: Qharabagh, Muhammad Fetrat, et al.
Pubblicazione: (2024)
Cross-Cultural Value Awareness in Large Vision-Language Models
di: Howard, Phillip, et al.
Pubblicazione: (2026)
di: Howard, Phillip, et al.
Pubblicazione: (2026)
Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models
di: Son, Jaemin, et al.
Pubblicazione: (2025)
di: Son, Jaemin, et al.
Pubblicazione: (2025)
Uncertainty Quantification for Multimodal Large Language Models with Incoherence-adjusted Semantic Volume
di: Lau, Gregory Kang Ruey, et al.
Pubblicazione: (2026)
di: Lau, Gregory Kang Ruey, et al.
Pubblicazione: (2026)
Bayesian Principles Improve Prompt Learning In Vision-Language Models
di: Kim, Mingyu, et al.
Pubblicazione: (2025)
di: Kim, Mingyu, et al.
Pubblicazione: (2025)
Towards Calibrated Robust Fine-Tuning of Vision-Language Models
di: Oh, Changdae, et al.
Pubblicazione: (2023)
di: Oh, Changdae, et al.
Pubblicazione: (2023)
Better Safe Than Sorry? Overreaction Problem of Vision Language Models in Visual Emergency Recognition
di: Choi, Dasol, et al.
Pubblicazione: (2025)
di: Choi, Dasol, et al.
Pubblicazione: (2025)
Focus Matters: Phase-Aware Suppression for Hallucination in Vision-Language Models
di: Kim, Sohyeon, et al.
Pubblicazione: (2026)
di: Kim, Sohyeon, et al.
Pubblicazione: (2026)
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
di: Oh, Youngtaek, et al.
Pubblicazione: (2024)
di: Oh, Youngtaek, et al.
Pubblicazione: (2024)
EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training
di: Du, Yiyang, et al.
Pubblicazione: (2026)
di: Du, Yiyang, et al.
Pubblicazione: (2026)
Evaluating Multimodal Generative AI with Korean Educational Standards
di: Park, Sanghee, et al.
Pubblicazione: (2025)
di: Park, Sanghee, et al.
Pubblicazione: (2025)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
di: Kim, Mingyeong, et al.
Pubblicazione: (2026)
di: Kim, Mingyeong, et al.
Pubblicazione: (2026)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
di: Ahn, Young Jin, et al.
Pubblicazione: (2024)
di: Ahn, Young Jin, et al.
Pubblicazione: (2024)
Hyperbolic Safety-Aware Vision-Language Models
di: Poppi, Tobia, et al.
Pubblicazione: (2025)
di: Poppi, Tobia, et al.
Pubblicazione: (2025)
Mind the Uncertainty in Human Disagreement: Evaluating Discrepancies between Model Predictions and Human Responses in VQA
di: Lan, Jian, et al.
Pubblicazione: (2024)
di: Lan, Jian, et al.
Pubblicazione: (2024)
MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
di: Hu, Wenbo, et al.
Pubblicazione: (2024)
di: Hu, Wenbo, et al.
Pubblicazione: (2024)
Navigating the Nuances: A Fine-grained Evaluation of Vision-Language Navigation
di: Wang, Zehao, et al.
Pubblicazione: (2024)
di: Wang, Zehao, et al.
Pubblicazione: (2024)
Can Vision Language Models Judge Action Quality? An Empirical Evaluation
di: Freitas, Miguel Monte e, et al.
Pubblicazione: (2026)
di: Freitas, Miguel Monte e, et al.
Pubblicazione: (2026)
ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
di: Zhang, Juntian, et al.
Pubblicazione: (2025)
di: Zhang, Juntian, et al.
Pubblicazione: (2025)
Documenti analoghi
-
GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
di: Park, Seongheon, et al.
Pubblicazione: (2025) -
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
di: Kundu, Souvik, et al.
Pubblicazione: (2025) -
Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding
di: Cho, Beomsik, et al.
Pubblicazione: (2025) -
Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
di: Padhi, Trilok, et al.
Pubblicazione: (2025) -
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
di: Lu, Yujie, et al.
Pubblicazione: (2024)