Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM Decoding
Fuente:
arXiv
Salvato in:
| Autori principali: | Mao, Shunqi, Zhang, Chaoyi, Cai, Weidong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Controllable Contextualized Image Captioning: Directing the Visual Narrative through User-Defined Highlights
di: Mao, Shunqi, et al.
Pubblicazione: (2024)
di: Mao, Shunqi, et al.
Pubblicazione: (2024)
DeepIcon: A Hierarchical Network for Layer-wise Icon Vectorization
di: Bing, Qi, et al.
Pubblicazione: (2024)
di: Bing, Qi, et al.
Pubblicazione: (2024)
Learning to Synthesize Graphics Programs for Geometric Artworks
di: Bing, Qi, et al.
Pubblicazione: (2024)
di: Bing, Qi, et al.
Pubblicazione: (2024)
Magnifier Prompt: Tackling Multimodal Hallucination via Extremely Simple Instructions
di: Fu, Yuhan, et al.
Pubblicazione: (2024)
di: Fu, Yuhan, et al.
Pubblicazione: (2024)
Exploring Annotation-free Image Captioning with Retrieval-augmented Pseudo Sentence Generation
di: Li, Zhiyuan, et al.
Pubblicazione: (2023)
di: Li, Zhiyuan, et al.
Pubblicazione: (2023)
Ctrl-Z Sampling: Scaling Diffusion Sampling with Controlled Random Zigzag Explorations
di: Mao, Shunqi, et al.
Pubblicazione: (2025)
di: Mao, Shunqi, et al.
Pubblicazione: (2025)
Enhancing Advanced Visual Reasoning Ability of Large Language Models
di: Li, Zhiyuan, et al.
Pubblicazione: (2024)
di: Li, Zhiyuan, et al.
Pubblicazione: (2024)
CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs
di: Kan, Zhehan, et al.
Pubblicazione: (2024)
di: Kan, Zhehan, et al.
Pubblicazione: (2024)
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
di: Chen, Xinlong, et al.
Pubblicazione: (2025)
di: Chen, Xinlong, et al.
Pubblicazione: (2025)
Multimodal Causal Reasoning Benchmark: Challenging Vision Large Language Models to Discern Causal Links Across Modalities
di: Li, Zhiyuan, et al.
Pubblicazione: (2024)
di: Li, Zhiyuan, et al.
Pubblicazione: (2024)
OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning
di: Li, Geng, et al.
Pubblicazione: (2026)
di: Li, Geng, et al.
Pubblicazione: (2026)
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
di: Ge, Haonan, et al.
Pubblicazione: (2025)
di: Ge, Haonan, et al.
Pubblicazione: (2025)
Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding
di: Suo, Wei, et al.
Pubblicazione: (2025)
di: Suo, Wei, et al.
Pubblicazione: (2025)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
di: Huang, Haiduo, et al.
Pubblicazione: (2025)
di: Huang, Haiduo, et al.
Pubblicazione: (2025)
Explore the Hallucination on Low-level Perception for MLLMs
di: Sun, Yinan, et al.
Pubblicazione: (2024)
di: Sun, Yinan, et al.
Pubblicazione: (2024)
See, Act, Adapt: Active Perception for Unsupervised Cross-Domain Visual Adaptation via Personalized VLM-Guided Agent
di: Tang, Tianci, et al.
Pubblicazione: (2026)
di: Tang, Tianci, et al.
Pubblicazione: (2026)
CHASD: Language Increment-Calibrated Contrastive Decoding against Hallucination in LVLMs
di: Huang, Xiaoyi, et al.
Pubblicazione: (2026)
di: Huang, Xiaoyi, et al.
Pubblicazione: (2026)
Adversarial Magnification to Deceive Deepfake Detection through Super Resolution
di: Coccomini, Davide Alessandro, et al.
Pubblicazione: (2024)
di: Coccomini, Davide Alessandro, et al.
Pubblicazione: (2024)
Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models
di: Huo, Fushuo, et al.
Pubblicazione: (2024)
di: Huo, Fushuo, et al.
Pubblicazione: (2024)
HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
An analysis of HOI: using a training-free method with multimodal visual foundation models when only the test set is available, without the training set
di: Ai, Chaoyi
Pubblicazione: (2024)
di: Ai, Chaoyi
Pubblicazione: (2024)
Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency
di: Wen, Ziqi, et al.
Pubblicazione: (2026)
di: Wen, Ziqi, et al.
Pubblicazione: (2026)
Deep Pulse-Signal Magnification for remote Heart Rate Estimation in Compressed Videos
di: Comas, Joaquim, et al.
Pubblicazione: (2024)
di: Comas, Joaquim, et al.
Pubblicazione: (2024)
Mitigating Hallucinations in Video Large Language Models via Spatiotemporal-Semantic Contrastive Decoding
di: Gao, Yuansheng, et al.
Pubblicazione: (2026)
di: Gao, Yuansheng, et al.
Pubblicazione: (2026)
Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding
di: Xu, Zhongxing, et al.
Pubblicazione: (2026)
di: Xu, Zhongxing, et al.
Pubblicazione: (2026)
GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
di: Wei, Tong, et al.
Pubblicazione: (2025)
di: Wei, Tong, et al.
Pubblicazione: (2025)
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning
di: Sun, Penglei, et al.
Pubblicazione: (2025)
di: Sun, Penglei, et al.
Pubblicazione: (2025)
Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI
di: Huang, Zheng, et al.
Pubblicazione: (2025)
di: Huang, Zheng, et al.
Pubblicazione: (2025)
Unifying VLM-Guided Flow Matching and Spectral Anomaly Detection for Interpretable Veterinary Diagnosis
di: Wang, Pu, et al.
Pubblicazione: (2026)
di: Wang, Pu, et al.
Pubblicazione: (2026)
ArtVLM: Attribute Recognition Through Vision-Based Prefix Language Modeling
di: Zhu, William Yicheng, et al.
Pubblicazione: (2024)
di: Zhu, William Yicheng, et al.
Pubblicazione: (2024)
Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models
di: Xing, Wenbin, et al.
Pubblicazione: (2026)
di: Xing, Wenbin, et al.
Pubblicazione: (2026)
BehaviorVLM: Unified Finetuning-Free Behavioral Understanding with Vision-Language Reasoning
di: Ke, Jingyang, et al.
Pubblicazione: (2026)
di: Ke, Jingyang, et al.
Pubblicazione: (2026)
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
di: Ji, Yicheng, et al.
Pubblicazione: (2025)
di: Ji, Yicheng, et al.
Pubblicazione: (2025)
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
di: Jiang, Yubo, et al.
Pubblicazione: (2026)
di: Jiang, Yubo, et al.
Pubblicazione: (2026)
SDCD: Structure-Disrupted Contrastive Decoding for Mitigating Hallucinations in Large Vision-Language Models
di: Xia, Yuxuan, et al.
Pubblicazione: (2026)
di: Xia, Yuxuan, et al.
Pubblicazione: (2026)
HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
di: Zhang, Jusheng, et al.
Pubblicazione: (2025)
di: Zhang, Jusheng, et al.
Pubblicazione: (2025)
Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual Guidance
di: Chen, Xinrong, et al.
Pubblicazione: (2026)
di: Chen, Xinrong, et al.
Pubblicazione: (2026)
The Collapse of Patches
di: Guo, Wei, et al.
Pubblicazione: (2025)
di: Guo, Wei, et al.
Pubblicazione: (2025)
CA-W3D: Leveraging Context-Aware Knowledge for Weakly Supervised Monocular 3D Detection
di: Liu, Chupeng, et al.
Pubblicazione: (2025)
di: Liu, Chupeng, et al.
Pubblicazione: (2025)
Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation
di: Jia, Sihang, et al.
Pubblicazione: (2026)
di: Jia, Sihang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Controllable Contextualized Image Captioning: Directing the Visual Narrative through User-Defined Highlights
di: Mao, Shunqi, et al.
Pubblicazione: (2024) -
DeepIcon: A Hierarchical Network for Layer-wise Icon Vectorization
di: Bing, Qi, et al.
Pubblicazione: (2024) -
Learning to Synthesize Graphics Programs for Geometric Artworks
di: Bing, Qi, et al.
Pubblicazione: (2024) -
Magnifier Prompt: Tackling Multimodal Hallucination via Extremely Simple Instructions
di: Fu, Yuhan, et al.
Pubblicazione: (2024) -
Exploring Annotation-free Image Captioning with Retrieval-augmented Pseudo Sentence Generation
di: Li, Zhiyuan, et al.
Pubblicazione: (2023)