Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding
Fuente:
arXiv
Salvato in:
| Autori principali: | Cho, Beomsik, Kim, Jaehyung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
What You See is What You Ask: Evaluating Audio Descriptions
di: Kala, Divy, et al.
Pubblicazione: (2025)
di: Kala, Divy, et al.
Pubblicazione: (2025)
VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation
di: Park, Seongheon, et al.
Pubblicazione: (2026)
di: Park, Seongheon, et al.
Pubblicazione: (2026)
On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models
di: Seo, Hoigi, et al.
Pubblicazione: (2025)
di: Seo, Hoigi, et al.
Pubblicazione: (2025)
Do You Keep an Eye on What I Ask? Mitigating Multimodal Hallucination via Attention-Guided Ensemble Decoding
di: Cho, Yeongjae, et al.
Pubblicazione: (2025)
di: Cho, Yeongjae, et al.
Pubblicazione: (2025)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
di: Kang, Seil, et al.
Pubblicazione: (2025)
di: Kang, Seil, et al.
Pubblicazione: (2025)
MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
di: Deng, Jingyuan, et al.
Pubblicazione: (2025)
di: Deng, Jingyuan, et al.
Pubblicazione: (2025)
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
di: Kim, Donghoon, et al.
Pubblicazione: (2025)
di: Kim, Donghoon, et al.
Pubblicazione: (2025)
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
di: Kundu, Souvik, et al.
Pubblicazione: (2025)
di: Kundu, Souvik, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
di: Min, Kyungmin, et al.
Pubblicazione: (2024)
di: Min, Kyungmin, et al.
Pubblicazione: (2024)
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
di: Deng, Ailin, et al.
Pubblicazione: (2024)
di: Deng, Ailin, et al.
Pubblicazione: (2024)
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
di: Ji, Yicheng, et al.
Pubblicazione: (2025)
di: Ji, Yicheng, et al.
Pubblicazione: (2025)
Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers
di: Wen, Yuxin, et al.
Pubblicazione: (2024)
di: Wen, Yuxin, et al.
Pubblicazione: (2024)
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
di: Lee, Yi-Lun, et al.
Pubblicazione: (2024)
di: Lee, Yi-Lun, et al.
Pubblicazione: (2024)
M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
di: Cho, Jaemin, et al.
Pubblicazione: (2024)
di: Cho, Jaemin, et al.
Pubblicazione: (2024)
BrainChat: Decoding Semantic Information from fMRI using Vision-language Pretrained Models
di: Huang, Wanaiu
Pubblicazione: (2024)
di: Huang, Wanaiu
Pubblicazione: (2024)
Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment
di: Xiao, Xin, et al.
Pubblicazione: (2024)
di: Xiao, Xin, et al.
Pubblicazione: (2024)
Analyzing The Language of Visual Tokens
di: Chan, David M., et al.
Pubblicazione: (2024)
di: Chan, David M., et al.
Pubblicazione: (2024)
How Much To Guide: Revisiting Adaptive Guidance in Classifier-Free Guidance Text-to-Vision Diffusion Models
di: Zhang, Huixuan, et al.
Pubblicazione: (2025)
di: Zhang, Huixuan, et al.
Pubblicazione: (2025)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
di: Tang, Zicong, et al.
Pubblicazione: (2025)
di: Tang, Zicong, et al.
Pubblicazione: (2025)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
di: Hwang, Yerin, et al.
Pubblicazione: (2025)
di: Hwang, Yerin, et al.
Pubblicazione: (2025)
Before Forgetting, Learn to Remember: Revisiting Foundational Learning Failures in LVLM Unlearning Benchmarks
di: Kwon, JuneHyoung, et al.
Pubblicazione: (2026)
di: Kwon, JuneHyoung, et al.
Pubblicazione: (2026)
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation
di: Yu, Seonghoon, et al.
Pubblicazione: (2026)
di: Yu, Seonghoon, et al.
Pubblicazione: (2026)
Revisiting the Role of Language Priors in Vision-Language Models
di: Lin, Zhiqiu, et al.
Pubblicazione: (2023)
di: Lin, Zhiqiu, et al.
Pubblicazione: (2023)
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
di: Oh, Youngtaek, et al.
Pubblicazione: (2024)
di: Oh, Youngtaek, et al.
Pubblicazione: (2024)
Seeing Culture: A Benchmark for Visual Reasoning and Grounding
di: Satar, Burak, et al.
Pubblicazione: (2025)
di: Satar, Burak, et al.
Pubblicazione: (2025)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
di: Ahn, Young Jin, et al.
Pubblicazione: (2024)
di: Ahn, Young Jin, et al.
Pubblicazione: (2024)
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
di: Chen, Xinlong, et al.
Pubblicazione: (2025)
di: Chen, Xinlong, et al.
Pubblicazione: (2025)
On the Limits of Token Reduction for Efficient Unified Vision Language Training
di: Chen, Siyi, et al.
Pubblicazione: (2026)
di: Chen, Siyi, et al.
Pubblicazione: (2026)
Brain-CLIPLM: Decoding Compressed Semantic Representations in EEG for Language Reconstruction
di: Yang, Xiaoli, et al.
Pubblicazione: (2026)
di: Yang, Xiaoli, et al.
Pubblicazione: (2026)
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
di: Li, Bangzheng, et al.
Pubblicazione: (2025)
di: Li, Bangzheng, et al.
Pubblicazione: (2025)
Be the Change You Want to See: Revisiting Remote Sensing Change Detection Practices
di: Rolih, Blaž, et al.
Pubblicazione: (2025)
di: Rolih, Blaž, et al.
Pubblicazione: (2025)
Seeing Symbols, Missing Cultures: Probing Vision-Language Models' Reasoning on Fire Imagery and Cultural Meaning
di: Yu, Haorui, et al.
Pubblicazione: (2025)
di: Yu, Haorui, et al.
Pubblicazione: (2025)
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models
di: He, Zoe Wanying, et al.
Pubblicazione: (2025)
di: He, Zoe Wanying, et al.
Pubblicazione: (2025)
Fwd2Bot: LVLM Visual Token Compression with Double Forward Bottleneck
di: Bulat, Adrian, et al.
Pubblicazione: (2025)
di: Bulat, Adrian, et al.
Pubblicazione: (2025)
LVLM-Aided Alignment of Task-Specific Vision Models
di: Koebler, Alexander, et al.
Pubblicazione: (2025)
di: Koebler, Alexander, et al.
Pubblicazione: (2025)
What You See is What You Classify: Black Box Attributions
di: Stalder, Steven, et al.
Pubblicazione: (2022)
di: Stalder, Steven, et al.
Pubblicazione: (2022)
Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning
di: Vaishnav, Mohit, et al.
Pubblicazione: (2026)
di: Vaishnav, Mohit, et al.
Pubblicazione: (2026)
OpenLVLM-MIA: A Controlled Benchmark Revealing the Limits of Membership Inference Attacks on Large Vision-Language Models
di: Miyamoto, Ryoto, et al.
Pubblicazione: (2025)
di: Miyamoto, Ryoto, et al.
Pubblicazione: (2025)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
di: Lin, Haokun, et al.
Pubblicazione: (2025)
di: Lin, Haokun, et al.
Pubblicazione: (2025)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
di: Zhang, Wanpeng, et al.
Pubblicazione: (2024)
di: Zhang, Wanpeng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
What You See is What You Ask: Evaluating Audio Descriptions
di: Kala, Divy, et al.
Pubblicazione: (2025) -
VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation
di: Park, Seongheon, et al.
Pubblicazione: (2026) -
On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models
di: Seo, Hoigi, et al.
Pubblicazione: (2025) -
Do You Keep an Eye on What I Ask? Mitigating Multimodal Hallucination via Attention-Guided Ensemble Decoding
di: Cho, Yeongjae, et al.
Pubblicazione: (2025) -
See What You Are Told: Visual Attention Sink in Large Multimodal Models
di: Kang, Seil, et al.
Pubblicazione: (2025)