Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Im, Eun Woo, Ali, Muhammad Kashif, Gupta, Vivek |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models
von: Huo, Fushuo, et al.
Veröffentlicht: (2024)
von: Huo, Fushuo, et al.
Veröffentlicht: (2024)
MAPWise: Evaluating Vision-Language Models for Advanced Map Queries
von: Mukhopadhyay, Srija, et al.
Veröffentlicht: (2024)
von: Mukhopadhyay, Srija, et al.
Veröffentlicht: (2024)
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
von: Chen, Xinlong, et al.
Veröffentlicht: (2025)
von: Chen, Xinlong, et al.
Veröffentlicht: (2025)
QAVA: Query-Agnostic Visual Attack to Large Vision-Language Models
von: Zhang, Yudong, et al.
Veröffentlicht: (2025)
von: Zhang, Yudong, et al.
Veröffentlicht: (2025)
Harnessing Meta-Learning for Controllable Full-Frame Video Stabilization
von: Ali, Muhammad Kashif, et al.
Veröffentlicht: (2025)
von: Ali, Muhammad Kashif, et al.
Veröffentlicht: (2025)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models
von: Shi, Yuheng, et al.
Veröffentlicht: (2026)
von: Shi, Yuheng, et al.
Veröffentlicht: (2026)
Harnessing Meta-Learning for Improving Full-Frame Video Stabilization
von: Ali, Muhammad Kashif, et al.
Veröffentlicht: (2024)
von: Ali, Muhammad Kashif, et al.
Veröffentlicht: (2024)
A-VL: Adaptive Attention for Large Vision-Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
SDCD: Structure-Disrupted Contrastive Decoding for Mitigating Hallucinations in Large Vision-Language Models
von: Xia, Yuxuan, et al.
Veröffentlicht: (2026)
von: Xia, Yuxuan, et al.
Veröffentlicht: (2026)
Deep Variational Bayesian Modeling of Haze Degradation Process
von: Im, Eun Woo, et al.
Veröffentlicht: (2024)
von: Im, Eun Woo, et al.
Veröffentlicht: (2024)
LVLM-COUNT: Enhancing the Counting Ability of Large Vision-Language Models
von: Qharabagh, Muhammad Fetrat, et al.
Veröffentlicht: (2024)
von: Qharabagh, Muhammad Fetrat, et al.
Veröffentlicht: (2024)
AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models
von: Rao, Zhifeng, et al.
Veröffentlicht: (2026)
von: Rao, Zhifeng, et al.
Veröffentlicht: (2026)
NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models
von: Pandya, Pranshu, et al.
Veröffentlicht: (2024)
von: Pandya, Pranshu, et al.
Veröffentlicht: (2024)
AgroGPT: Efficient Agricultural Vision-Language Model with Expert Tuning
von: Awais, Muhammad, et al.
Veröffentlicht: (2024)
von: Awais, Muhammad, et al.
Veröffentlicht: (2024)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
von: Min, Kyungmin, et al.
Veröffentlicht: (2024)
von: Min, Kyungmin, et al.
Veröffentlicht: (2024)
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
von: Lee, Yi-Lun, et al.
Veröffentlicht: (2024)
von: Lee, Yi-Lun, et al.
Veröffentlicht: (2024)
Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models
von: Wang, Zhiqiang, et al.
Veröffentlicht: (2026)
von: Wang, Zhiqiang, et al.
Veröffentlicht: (2026)
Decoding Neighborhood Environments with Large Language Models
von: Cart, Andrew, et al.
Veröffentlicht: (2025)
von: Cart, Andrew, et al.
Veröffentlicht: (2025)
Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual Guidance
von: Chen, Xinrong, et al.
Veröffentlicht: (2026)
von: Chen, Xinrong, et al.
Veröffentlicht: (2026)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)
von: Manevich, Avshalom, et al.
Veröffentlicht: (2024)
von: Manevich, Avshalom, et al.
Veröffentlicht: (2024)
MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models
von: Paik, Gio, et al.
Veröffentlicht: (2025)
von: Paik, Gio, et al.
Veröffentlicht: (2025)
Speculative Decoding Reimagined for Multimodal Large Language Models
von: Lin, Luxi, et al.
Veröffentlicht: (2025)
von: Lin, Luxi, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
von: Wang, Xintong, et al.
Veröffentlicht: (2024)
von: Wang, Xintong, et al.
Veröffentlicht: (2024)
CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models
von: Cao, Zongsheng, et al.
Veröffentlicht: (2025)
von: Cao, Zongsheng, et al.
Veröffentlicht: (2025)
Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models
von: Zou, Zhengtao, et al.
Veröffentlicht: (2025)
von: Zou, Zhengtao, et al.
Veröffentlicht: (2025)
A Visual Semantic Adaptive Watermark grounded by Prefix-Tuning for Large Vision-Language Model
von: Zheng, Qi, et al.
Veröffentlicht: (2026)
von: Zheng, Qi, et al.
Veröffentlicht: (2026)
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models
von: Choi, Dasol, et al.
Veröffentlicht: (2026)
von: Choi, Dasol, et al.
Veröffentlicht: (2026)
Phantasia: Context-Adaptive Backdoors in Vision Language Models
von: Tran, Nam Duong, et al.
Veröffentlicht: (2026)
von: Tran, Nam Duong, et al.
Veröffentlicht: (2026)
Freeze and Reveal: Exposing Modality Bias in Vision-Language Models
von: Kavuri, Vivek Hruday, et al.
Veröffentlicht: (2025)
von: Kavuri, Vivek Hruday, et al.
Veröffentlicht: (2025)
VLAgeBench: Benchmarking Large Vision-Language Models for Zero-Shot Human Age Estimation
von: Sajib, Rakib Hossain, et al.
Veröffentlicht: (2026)
von: Sajib, Rakib Hossain, et al.
Veröffentlicht: (2026)
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
von: Li, Chaoyu, et al.
Veröffentlicht: (2024)
von: Li, Chaoyu, et al.
Veröffentlicht: (2024)
Reflexive Guidance: Improving OoDD in Vision-Language Models via Self-Guided Image-Adaptive Concept Generation
von: Kim, Jihyo, et al.
Veröffentlicht: (2024)
von: Kim, Jihyo, et al.
Veröffentlicht: (2024)
AT-SNN: Adaptive Tokens for Vision Transformer on Spiking Neural Network
von: Kang, Donghwa, et al.
Veröffentlicht: (2024)
von: Kang, Donghwa, et al.
Veröffentlicht: (2024)
From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models
von: Lee, Jihoon, et al.
Veröffentlicht: (2025)
von: Lee, Jihoon, et al.
Veröffentlicht: (2025)
Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization
von: Jia, Chenwei, et al.
Veröffentlicht: (2026)
von: Jia, Chenwei, et al.
Veröffentlicht: (2026)
Understanding Counting Mechanisms in Large Language and Vision-Language Models
von: Hasani, Hosein, et al.
Veröffentlicht: (2025)
von: Hasani, Hosein, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models
von: Huo, Fushuo, et al.
Veröffentlicht: (2024) -
MAPWise: Evaluating Vision-Language Models for Advanced Map Queries
von: Mukhopadhyay, Srija, et al.
Veröffentlicht: (2024) -
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
von: Chen, Xinlong, et al.
Veröffentlicht: (2025) -
QAVA: Query-Agnostic Visual Attack to Large Vision-Language Models
von: Zhang, Yudong, et al.
Veröffentlicht: (2025) -
Harnessing Meta-Learning for Controllable Full-Frame Video Stabilization
von: Ali, Muhammad Kashif, et al.
Veröffentlicht: (2025)