IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhu, Lanyun, Ji, Deyi, Chen, Tianrun, Xu, Peng, Ye, Jieping, Liu, Jun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLaFS: When Large Language Models Meet Few-Shot Segmentation
por: Zhu, Lanyun, et al.
Publicado: (2023)
por: Zhu, Lanyun, et al.
Publicado: (2023)
Not Every Patch is Needed: Towards a More Efficient and Effective Backbone for Video-based Person Re-identification
por: Zhu, Lanyun, et al.
Publicado: (2025)
por: Zhu, Lanyun, et al.
Publicado: (2025)
Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
por: Zhu, Lanyun, et al.
Publicado: (2025)
por: Zhu, Lanyun, et al.
Publicado: (2025)
Discrete Latent Perspective Learning for Segmentation and Detection
por: Ji, Deyi, et al.
Publicado: (2024)
por: Ji, Deyi, et al.
Publicado: (2024)
Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models
por: Huo, Fushuo, et al.
Publicado: (2024)
por: Huo, Fushuo, et al.
Publicado: (2024)
xLSTM-UNet can be an Effective 2D & 3D Medical Image Segmentation Backbone with Vision-LSTM (ViL) better than its Mamba Counterpart
por: Chen, Tianrun, et al.
Publicado: (2024)
por: Chen, Tianrun, et al.
Publicado: (2024)
POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation
por: Zhu, Lanyun, et al.
Publicado: (2025)
por: Zhu, Lanyun, et al.
Publicado: (2025)
Reasoning3D -- Grounding and Reasoning in 3D: Fine-Grained Zero-Shot Open-Vocabulary 3D Reasoning Part Segmentation via Large Vision-Language Models
por: Chen, Tianrun, et al.
Publicado: (2024)
por: Chen, Tianrun, et al.
Publicado: (2024)
Breaking the Box: Enhancing Remote Sensing Image Segmentation with Freehand Sketches
por: Zang, Ying, et al.
Publicado: (2025)
por: Zang, Ying, et al.
Publicado: (2025)
StreamSense: Streaming Social Task Detection with Selective Vision-Language Model Routing
por: Wang, Han, et al.
Publicado: (2026)
por: Wang, Han, et al.
Publicado: (2026)
Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization
por: Lyu, Xinyu, et al.
Publicado: (2024)
por: Lyu, Xinyu, et al.
Publicado: (2024)
SAM3-Adapter: Efficient Adaptation of Segment Anything 3 for Camouflage Object Segmentation, Shadow Detection, and Medical Image Segmentation
por: Chen, Tianrun, et al.
Publicado: (2025)
por: Chen, Tianrun, et al.
Publicado: (2025)
StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression
por: Liu, Xuanyi, et al.
Publicado: (2026)
por: Liu, Xuanyi, et al.
Publicado: (2026)
View-Centric Multi-Object Tracking with Homographic Matching in Moving UAV
por: Ji, Deyi, et al.
Publicado: (2024)
por: Ji, Deyi, et al.
Publicado: (2024)
From Air to Wear: Personalized 3D Digital Fashion with AR/VR Immersive 3D Sketching
por: Zang, Ying, et al.
Publicado: (2025)
por: Zang, Ying, et al.
Publicado: (2025)
SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More
por: Chen, Tianrun, et al.
Publicado: (2024)
por: Chen, Tianrun, et al.
Publicado: (2024)
Let Human Sketches Help: Empowering Challenging Image Segmentation Task with Freehand Sketches
por: Zang, Ying, et al.
Publicado: (2025)
por: Zang, Ying, et al.
Publicado: (2025)
YARD: Y-Architecture Register Decoding for Efficient Hallucination Mitigation in Large Vision-Language Models
por: Chen, Ting, et al.
Publicado: (2026)
por: Chen, Ting, et al.
Publicado: (2026)
Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding
por: Suo, Wei, et al.
Publicado: (2025)
por: Suo, Wei, et al.
Publicado: (2025)
Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors
por: Zang, Ying, et al.
Publicado: (2026)
por: Zang, Ying, et al.
Publicado: (2026)
Structural and Statistical Texture Knowledge Distillation and Learning for Segmentation
por: Ji, Deyi, et al.
Publicado: (2025)
por: Ji, Deyi, et al.
Publicado: (2025)
Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation
por: Qu, Xiaoye, et al.
Publicado: (2024)
por: Qu, Xiaoye, et al.
Publicado: (2024)
SDCD: Structure-Disrupted Contrastive Decoding for Mitigating Hallucinations in Large Vision-Language Models
por: Xia, Yuxuan, et al.
Publicado: (2026)
por: Xia, Yuxuan, et al.
Publicado: (2026)
AddressCLIP: Empowering Vision-Language Models for City-wide Image Address Localization
por: Xu, Shixiong, et al.
Publicado: (2024)
por: Xu, Shixiong, et al.
Publicado: (2024)
CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models
por: Cao, Zongsheng, et al.
Publicado: (2025)
por: Cao, Zongsheng, et al.
Publicado: (2025)
Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language Models
por: Wu, Mingrui, et al.
Publicado: (2024)
por: Wu, Mingrui, et al.
Publicado: (2024)
SAKED: Mitigating Hallucination in Large Vision-Language Models via Stability-Aware Knowledge Enhanced Decoding
por: Li, Zhaoxu, et al.
Publicado: (2026)
por: Li, Zhaoxu, et al.
Publicado: (2026)
Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual Guidance
por: Chen, Xinrong, et al.
Publicado: (2026)
por: Chen, Xinrong, et al.
Publicado: (2026)
4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation
por: Zang, Ying, et al.
Publicado: (2026)
por: Zang, Ying, et al.
Publicado: (2026)
Mitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration
por: Zhu, Younan, et al.
Publicado: (2025)
por: Zhu, Younan, et al.
Publicado: (2025)
Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
por: Jiang, Chaoya, et al.
Publicado: (2023)
por: Jiang, Chaoya, et al.
Publicado: (2023)
Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
por: Zhang, Xiaofeng, et al.
Publicado: (2024)
por: Zhang, Xiaofeng, et al.
Publicado: (2024)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
por: Min, Kyungmin, et al.
Publicado: (2024)
por: Min, Kyungmin, et al.
Publicado: (2024)
SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding
por: Park, Woohyeon, et al.
Publicado: (2025)
por: Park, Woohyeon, et al.
Publicado: (2025)
HD-VGGT: High-Resolution Visual Geometry Transformer
por: Chen, Tianrun, et al.
Publicado: (2026)
por: Chen, Tianrun, et al.
Publicado: (2026)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
por: Wang, Zihu, et al.
Publicado: (2025)
por: Wang, Zihu, et al.
Publicado: (2025)
Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens
por: Jiang, Zhangqi, et al.
Publicado: (2024)
por: Jiang, Zhangqi, et al.
Publicado: (2024)
Mitigating Hallucination in Visual-Language Models via Re-Balancing Contrastive Decoding
por: Liang, Xiaoyu, et al.
Publicado: (2024)
por: Liang, Xiaoyu, et al.
Publicado: (2024)
Mitigating Hallucinations in Large Vision-Language Models via Causal Route Gating
por: Cheng, Zhe, et al.
Publicado: (2026)
por: Cheng, Zhe, et al.
Publicado: (2026)
Mitigating Hallucinations in Video Large Language Models via Spatiotemporal-Semantic Contrastive Decoding
por: Gao, Yuansheng, et al.
Publicado: (2026)
por: Gao, Yuansheng, et al.
Publicado: (2026)
Ejemplares similares
-
LLaFS: When Large Language Models Meet Few-Shot Segmentation
por: Zhu, Lanyun, et al.
Publicado: (2023) -
Not Every Patch is Needed: Towards a More Efficient and Effective Backbone for Video-based Person Re-identification
por: Zhu, Lanyun, et al.
Publicado: (2025) -
Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
por: Zhu, Lanyun, et al.
Publicado: (2025) -
Discrete Latent Perspective Learning for Segmentation and Detection
por: Ji, Deyi, et al.
Publicado: (2024) -
Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models
por: Huo, Fushuo, et al.
Publicado: (2024)