SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, Yiyang, Shi, Liang, Zhang, Yitian, Xu, Yi, Fu, Yun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Distorted or Fabricated? A Survey on Hallucination in Video LLMs
por: Huang, Yiyang, et al.
Publicado: (2026)
por: Huang, Yiyang, et al.
Publicado: (2026)
Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation
por: Cheng, Ruoxi, et al.
Publicado: (2026)
por: Cheng, Ruoxi, et al.
Publicado: (2026)
Cross-Modal Attention Calibration for LVLM Hallucination Mitigation
por: Li, Jiaming, et al.
Publicado: (2025)
por: Li, Jiaming, et al.
Publicado: (2025)
Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation
por: Mao, Jiawei, et al.
Publicado: (2026)
por: Mao, Jiawei, et al.
Publicado: (2026)
Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?
por: Feng, Mingqian, et al.
Publicado: (2024)
por: Feng, Mingqian, et al.
Publicado: (2024)
D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
por: Huang, Yiyang, et al.
Publicado: (2025)
por: Huang, Yiyang, et al.
Publicado: (2025)
Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats
por: Qian, Jiaye, et al.
Publicado: (2025)
por: Qian, Jiaye, et al.
Publicado: (2025)
Accessing Vision Foundation Models via ImageNet-1K
por: Zhang, Yitian, et al.
Publicado: (2024)
por: Zhang, Yitian, et al.
Publicado: (2024)
MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
por: Deng, Jingyuan, et al.
Publicado: (2025)
por: Deng, Jingyuan, et al.
Publicado: (2025)
Fighting Hallucinations with Counterfactuals: Diffusion-Guided Perturbations for LVLM Hallucination Suppression
por: Dastmalchi, Hamidreza, et al.
Publicado: (2026)
por: Dastmalchi, Hamidreza, et al.
Publicado: (2026)
Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward
por: Fan, Zhiyuan, et al.
Publicado: (2025)
por: Fan, Zhiyuan, et al.
Publicado: (2025)
Before Forgetting, Learn to Remember: Revisiting Foundational Learning Failures in LVLM Unlearning Benchmarks
por: Kwon, JuneHyoung, et al.
Publicado: (2026)
por: Kwon, JuneHyoung, et al.
Publicado: (2026)
TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models
por: Shao, Wenqi, et al.
Publicado: (2023)
por: Shao, Wenqi, et al.
Publicado: (2023)
ClueTracer: Question-to-Vision Clue Tracing for Training-Free Hallucination Suppression in Multimodal Reasoning
por: Xi, Gongli, et al.
Publicado: (2026)
por: Xi, Gongli, et al.
Publicado: (2026)
LVLM-Aided Alignment of Task-Specific Vision Models
por: Koebler, Alexander, et al.
Publicado: (2025)
por: Koebler, Alexander, et al.
Publicado: (2025)
Adversarial Defense in Vision-Language Models: An Overview
por: Fu, Xiaowei, et al.
Publicado: (2026)
por: Fu, Xiaowei, et al.
Publicado: (2026)
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models
por: Jin, Haibo, et al.
Publicado: (2025)
por: Jin, Haibo, et al.
Publicado: (2025)
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge
por: Zhao, Yaqi, et al.
Publicado: (2024)
por: Zhao, Yaqi, et al.
Publicado: (2024)
General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting
por: Lange, Bernard, et al.
Publicado: (2025)
por: Lange, Bernard, et al.
Publicado: (2025)
OmniCT: Towards a Unified Slice-Volume LVLM for Comprehensive CT Analysis
por: Lin, Tianwei, et al.
Publicado: (2026)
por: Lin, Tianwei, et al.
Publicado: (2026)
LVLM-COUNT: Enhancing the Counting Ability of Large Vision-Language Models
por: Qharabagh, Muhammad Fetrat, et al.
Publicado: (2024)
por: Qharabagh, Muhammad Fetrat, et al.
Publicado: (2024)
Don't Judge by the Look: Towards Motion Coherent Video Representation
por: Zhang, Yitian, et al.
Publicado: (2024)
por: Zhang, Yitian, et al.
Publicado: (2024)
Security Tensors as a Cross-Modal Bridge: Extending Text-Aligned Safety to Vision in LVLM
por: Li, Shen, et al.
Publicado: (2025)
por: Li, Shen, et al.
Publicado: (2025)
EmoFeedback$^2$: Reinforcement of Continuous Emotional Image Generation via LVLM-based Reward and Textual Feedback
por: Jia, Jingyang, et al.
Publicado: (2025)
por: Jia, Jingyang, et al.
Publicado: (2025)
Gestura: A LVLM-Powered System Bridging Motion and Semantics for Real-Time Free-Form Gesture Understanding
por: Li, Zhuoming, et al.
Publicado: (2025)
por: Li, Zhuoming, et al.
Publicado: (2025)
MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
por: Liu, Yi, et al.
Publicado: (2025)
por: Liu, Yi, et al.
Publicado: (2025)
Focus Matters: Phase-Aware Suppression for Hallucination in Vision-Language Models
por: Kim, Sohyeon, et al.
Publicado: (2026)
por: Kim, Sohyeon, et al.
Publicado: (2026)
Geo-LLaVA: A Large Multi-Modal Model for Solving Geometry Math Problems with Meta In-Context Learning
por: Xu, Shihao, et al.
Publicado: (2024)
por: Xu, Shihao, et al.
Publicado: (2024)
VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation
por: Park, Seongheon, et al.
Publicado: (2026)
por: Park, Seongheon, et al.
Publicado: (2026)
Locating Demographic Bias at the Attention-Head Level in CLIP's Vision Encoder
por: Yasser, Alaa, et al.
Publicado: (2026)
por: Yasser, Alaa, et al.
Publicado: (2026)
Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding
por: Suo, Wei, et al.
Publicado: (2025)
por: Suo, Wei, et al.
Publicado: (2025)
NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors
por: Ren, Lingfeng, et al.
Publicado: (2026)
por: Ren, Lingfeng, et al.
Publicado: (2026)
MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
por: Li, Hui, et al.
Publicado: (2025)
por: Li, Hui, et al.
Publicado: (2025)
Suppress Content Shift: Better Diffusion Features via Off-the-Shelf Generation Techniques
por: Meng, Benyuan, et al.
Publicado: (2024)
por: Meng, Benyuan, et al.
Publicado: (2024)
Agglomerating Large Vision Encoders via Distillation for VFSS Segmentation
por: Zeng, Chengxi, et al.
Publicado: (2025)
por: Zeng, Chengxi, et al.
Publicado: (2025)
TARAC: Mitigating Hallucination in LVLMs via Temporal Attention Real-time Accumulative Connection
por: Jiang, Lei, et al.
Publicado: (2025)
por: Jiang, Lei, et al.
Publicado: (2025)
Learning Fine-grained Domain Generalization via Hyperbolic State Space Hallucination
por: Bi, Qi, et al.
Publicado: (2025)
por: Bi, Qi, et al.
Publicado: (2025)
ReflexFlow: Rethinking Learning Objective for Exposure Bias Alleviation in Flow Matching
por: Huang, Guanbo, et al.
Publicado: (2025)
por: Huang, Guanbo, et al.
Publicado: (2025)
OpenBias: Open-set Bias Detection in Text-to-Image Generative Models
por: D'Incà, Moreno, et al.
Publicado: (2024)
por: D'Incà, Moreno, et al.
Publicado: (2024)
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
por: Li, Shuo, et al.
Publicado: (2025)
por: Li, Shuo, et al.
Publicado: (2025)
Ejemplares similares
-
Distorted or Fabricated? A Survey on Hallucination in Video LLMs
por: Huang, Yiyang, et al.
Publicado: (2026) -
Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation
por: Cheng, Ruoxi, et al.
Publicado: (2026) -
Cross-Modal Attention Calibration for LVLM Hallucination Mitigation
por: Li, Jiaming, et al.
Publicado: (2025) -
Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation
por: Mao, Jiawei, et al.
Publicado: (2026) -
Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?
por: Feng, Mingqian, et al.
Publicado: (2024)