Context-Aware Decoding for Faithful Vision-Language Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Fazli, Mehrdad, Wei, Bowen, Zhu, Ziwei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
di: Fazli, Mehrdad, et al.
Pubblicazione: (2025)
di: Fazli, Mehrdad, et al.
Pubblicazione: (2025)
Long Context Transfer from Language to Vision
di: Zhang, Peiyuan, et al.
Pubblicazione: (2024)
di: Zhang, Peiyuan, et al.
Pubblicazione: (2024)
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
di: Kang, Jialiang, et al.
Pubblicazione: (2025)
di: Kang, Jialiang, et al.
Pubblicazione: (2025)
Context-Aware Autoregressive Models for Multi-Conditional Image Generation
di: Chen, Yixiao, et al.
Pubblicazione: (2025)
di: Chen, Yixiao, et al.
Pubblicazione: (2025)
On the Faithfulness of Vision Transformer Explanations
di: Wu, Junyi, et al.
Pubblicazione: (2024)
di: Wu, Junyi, et al.
Pubblicazione: (2024)
Vision-Language Binding in In-Context Image Generation
di: Ge, Chris, et al.
Pubblicazione: (2026)
di: Ge, Chris, et al.
Pubblicazione: (2026)
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
di: Kulkarni, Yogesh, et al.
Pubblicazione: (2024)
di: Kulkarni, Yogesh, et al.
Pubblicazione: (2024)
VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
di: Zheng, Dian, et al.
Pubblicazione: (2025)
di: Zheng, Dian, et al.
Pubblicazione: (2025)
FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models
di: Jing, Liqiang, et al.
Pubblicazione: (2023)
di: Jing, Liqiang, et al.
Pubblicazione: (2023)
Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning
di: Liu, Shih-Wen, et al.
Pubblicazione: (2025)
di: Liu, Shih-Wen, et al.
Pubblicazione: (2025)
Bridging Vision and Language for Robust Context-Aware Surgical Point Tracking: The VL-SurgPT Dataset and Benchmark
di: Zhou, Rulin, et al.
Pubblicazione: (2025)
di: Zhou, Rulin, et al.
Pubblicazione: (2025)
Stencil: Subject-Driven Generation with Context Guidance
di: Chen, Gordon, et al.
Pubblicazione: (2025)
di: Chen, Gordon, et al.
Pubblicazione: (2025)
Dude: Dual Distribution-Aware Context Prompt Learning For Large Vision-Language Model
di: Nguyen, Duy M. H., et al.
Pubblicazione: (2024)
di: Nguyen, Duy M. H., et al.
Pubblicazione: (2024)
SAKED: Mitigating Hallucination in Large Vision-Language Models via Stability-Aware Knowledge Enhanced Decoding
di: Li, Zhaoxu, et al.
Pubblicazione: (2026)
di: Li, Zhaoxu, et al.
Pubblicazione: (2026)
Context Diffusion: In-Context Aware Image Generation
di: Najdenkoska, Ivona, et al.
Pubblicazione: (2023)
di: Najdenkoska, Ivona, et al.
Pubblicazione: (2023)
Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context
di: Wang, Zhaowei, et al.
Pubblicazione: (2026)
di: Wang, Zhaowei, et al.
Pubblicazione: (2026)
Towards Multimodal In-Context Learning for Vision & Language Models
di: Doveh, Sivan, et al.
Pubblicazione: (2024)
di: Doveh, Sivan, et al.
Pubblicazione: (2024)
Evaluating Reasoning Faithfulness in Medical Vision-Language Models using Multimodal Perturbations
di: Moll, Johannes, et al.
Pubblicazione: (2025)
di: Moll, Johannes, et al.
Pubblicazione: (2025)
VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models
di: Qiu, Haoyi, et al.
Pubblicazione: (2024)
di: Qiu, Haoyi, et al.
Pubblicazione: (2024)
Large Vision-Language Models as Emotion Recognizers in Context Awareness
di: Lei, Yuxuan, et al.
Pubblicazione: (2024)
di: Lei, Yuxuan, et al.
Pubblicazione: (2024)
Branch, or Layer? Zeroth-Order Optimization for Continual Learning of Vision-Language Models
di: Liu, Ziwei, et al.
Pubblicazione: (2025)
di: Liu, Ziwei, et al.
Pubblicazione: (2025)
Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models
di: Qi, Jianing, et al.
Pubblicazione: (2025)
di: Qi, Jianing, et al.
Pubblicazione: (2025)
MMSpec: Benchmarking Speculative Decoding for Vision-Language Models
di: Shen, Hui, et al.
Pubblicazione: (2026)
di: Shen, Hui, et al.
Pubblicazione: (2026)
IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding
di: Zhu, Lanyun, et al.
Pubblicazione: (2024)
di: Zhu, Lanyun, et al.
Pubblicazione: (2024)
ChartQA-X: Generating Explanations for Visual Chart Reasoning
di: Hegde, Shamanthak, et al.
Pubblicazione: (2025)
di: Hegde, Shamanthak, et al.
Pubblicazione: (2025)
Negation-Aware Test-Time Adaptation for Vision-Language Models
di: Han, Haochen, et al.
Pubblicazione: (2025)
di: Han, Haochen, et al.
Pubblicazione: (2025)
CASCADE: Context-Aware Relaxation for Speculative Image Decoding
di: Yildirim, Selin, et al.
Pubblicazione: (2026)
di: Yildirim, Selin, et al.
Pubblicazione: (2026)
Context-Aware Token Selection and Packing for Enhanced Vision Transformer
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations
di: Ding, Sihao, et al.
Pubblicazione: (2025)
di: Ding, Sihao, et al.
Pubblicazione: (2025)
Optimizing Vision-Language Interactions Through Decoder-Only Models
di: Tanaka, Kaito, et al.
Pubblicazione: (2024)
di: Tanaka, Kaito, et al.
Pubblicazione: (2024)
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
di: Yang, Fan, et al.
Pubblicazione: (2026)
di: Yang, Fan, et al.
Pubblicazione: (2026)
Dynamic Context-Aware Scene Reasoning Using Vision-Language Alignment in Zero-Shot Real-World Scenarios
di: Rajiv, Manjunath Prasad Holenarasipura, et al.
Pubblicazione: (2025)
di: Rajiv, Manjunath Prasad Holenarasipura, et al.
Pubblicazione: (2025)
Efficient and Context-Aware Label Propagation for Zero-/Few-Shot Training-Free Adaptation of Vision-Language Model
di: Li, Yushu, et al.
Pubblicazione: (2024)
di: Li, Yushu, et al.
Pubblicazione: (2024)
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
di: Li, Chaoyu, et al.
Pubblicazione: (2024)
di: Li, Chaoyu, et al.
Pubblicazione: (2024)
Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
di: Zhang, Ce, et al.
Pubblicazione: (2025)
di: Zhang, Ce, et al.
Pubblicazione: (2025)
IIR-VLM: In-Context Instance-level Recognition for Large Vision-Language Models
di: Shi, Liang, et al.
Pubblicazione: (2026)
di: Shi, Liang, et al.
Pubblicazione: (2026)
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
di: Rahman, Md Ashikur, et al.
Pubblicazione: (2026)
di: Rahman, Md Ashikur, et al.
Pubblicazione: (2026)
FRAP: Faithful and Realistic Text-to-Image Generation with Adaptive Prompt Weighting
di: Jiang, Liyao, et al.
Pubblicazione: (2024)
di: Jiang, Liyao, et al.
Pubblicazione: (2024)
Modality-Agnostic fMRI Decoding of Vision and Language
di: Nikolaus, Mitja, et al.
Pubblicazione: (2024)
di: Nikolaus, Mitja, et al.
Pubblicazione: (2024)
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
di: Kulkarni, Yogesh, et al.
Pubblicazione: (2025)
di: Kulkarni, Yogesh, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
di: Fazli, Mehrdad, et al.
Pubblicazione: (2025) -
Long Context Transfer from Language to Vision
di: Zhang, Peiyuan, et al.
Pubblicazione: (2024) -
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
di: Kang, Jialiang, et al.
Pubblicazione: (2025) -
Context-Aware Autoregressive Models for Multi-Conditional Image Generation
di: Chen, Yixiao, et al.
Pubblicazione: (2025) -
On the Faithfulness of Vision Transformer Explanations
di: Wu, Junyi, et al.
Pubblicazione: (2024)