Mitigating Cross-Image Information Leakage in LVLMs for Multi-Image Tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Park, Yeji, Lee, Minyoung, Chun, Sanghyuk, Choe, Junsuk |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
por: Park, Yeji, et al.
Publicado: (2024)
por: Park, Yeji, et al.
Publicado: (2024)
Enhancing Multi-Image Understanding through Delimiter Token Scaling
por: Lee, Minyoung, et al.
Publicado: (2026)
por: Lee, Minyoung, et al.
Publicado: (2026)
LMLT: Low-to-high Multi-Level Vision Transformer for Image Super-Resolution
por: Kim, Jeongsoo, et al.
Publicado: (2024)
por: Kim, Jeongsoo, et al.
Publicado: (2024)
Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs
por: Lee, Insu, et al.
Publicado: (2025)
por: Lee, Insu, et al.
Publicado: (2025)
Sampling Bag of Views for Open-Vocabulary Object Detection
por: Choi, Hojun, et al.
Publicado: (2024)
por: Choi, Hojun, et al.
Publicado: (2024)
RoCOCO: Robustness Benchmark of MS-COCO to Stress-test Image-Text Matching Models
por: Park, Seulki, et al.
Publicado: (2023)
por: Park, Seulki, et al.
Publicado: (2023)
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
por: Ge, Haonan, et al.
Publicado: (2025)
por: Ge, Haonan, et al.
Publicado: (2025)
What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
por: Kang, Inha, et al.
Publicado: (2025)
por: Kang, Inha, et al.
Publicado: (2025)
Optimizing LVLMs with On-Policy Data for Effective Hallucination Mitigation
por: Yu, Chengzhi, et al.
Publicado: (2025)
por: Yu, Chengzhi, et al.
Publicado: (2025)
Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification
por: Sun, Han, et al.
Publicado: (2026)
por: Sun, Han, et al.
Publicado: (2026)
From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
por: Zhang, Xiaofeng, et al.
Publicado: (2024)
por: Zhang, Xiaofeng, et al.
Publicado: (2024)
Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation
por: Yang, Xiaochen, et al.
Publicado: (2026)
por: Yang, Xiaochen, et al.
Publicado: (2026)
VidLBEval: Benchmarking and Mitigating Language Bias in Video-Involved LVLMs
por: Yang, Yiming, et al.
Publicado: (2025)
por: Yang, Yiming, et al.
Publicado: (2025)
Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model
por: Zhao, Min, et al.
Publicado: (2024)
por: Zhao, Min, et al.
Publicado: (2024)
Causally-Grounded Dual-Path Attention Intervention for Object Hallucination Mitigation in LVLMs
por: Yu, Liu, et al.
Publicado: (2025)
por: Yu, Liu, et al.
Publicado: (2025)
MHSA: A Lightweight Framework for Mitigating Hallucinations via Steered Attention in LVLMs
por: Ding, Wei, et al.
Publicado: (2026)
por: Ding, Wei, et al.
Publicado: (2026)
CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs
por: Kan, Zhehan, et al.
Publicado: (2024)
por: Kan, Zhehan, et al.
Publicado: (2024)
DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
por: Ventura, Mor, et al.
Publicado: (2025)
por: Ventura, Mor, et al.
Publicado: (2025)
TARAC: Mitigating Hallucination in LVLMs via Temporal Attention Real-time Accumulative Connection
por: Jiang, Lei, et al.
Publicado: (2025)
por: Jiang, Lei, et al.
Publicado: (2025)
MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
por: Liu, Ziyu, et al.
Publicado: (2024)
por: Liu, Ziyu, et al.
Publicado: (2024)
OVS Meets Continual Learning: Towards Sustainable Open-Vocabulary Segmentation
por: Hwang, Dongjun, et al.
Publicado: (2024)
por: Hwang, Dongjun, et al.
Publicado: (2024)
Revealing Multi-View Hallucination in Large Vision-Language Models
por: Park, Wooje, et al.
Publicado: (2026)
por: Park, Wooje, et al.
Publicado: (2026)
See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis
por: Park, Jaehyun, et al.
Publicado: (2026)
por: Park, Jaehyun, et al.
Publicado: (2026)
Improved Probabilistic Image-Text Representations
por: Chun, Sanghyuk
Publicado: (2023)
por: Chun, Sanghyuk
Publicado: (2023)
Attention-space Contrastive Guidance for Efficient Hallucination Mitigation in LVLMs
por: Jo, Yujin, et al.
Publicado: (2026)
por: Jo, Yujin, et al.
Publicado: (2026)
When RAG Hurts: Diagnosing and Mitigating Attention Distraction in Retrieval-Augmented LVLMs
por: Zhao, Beidi, et al.
Publicado: (2026)
por: Zhao, Beidi, et al.
Publicado: (2026)
CrossMed: A Multimodal Cross-Task Benchmark for Compositional Generalization in Medical Imaging
por: Singh, Pooja, et al.
Publicado: (2025)
por: Singh, Pooja, et al.
Publicado: (2025)
Countering the Over-Reliance Trap: Mitigating Object Hallucination for LVLMs via a Self-Validation Framework
por: Liu, Shiyu, et al.
Publicado: (2026)
por: Liu, Shiyu, et al.
Publicado: (2026)
Reciprocal Attention Mixing Transformer for Lightweight Image Restoration
por: Choi, Haram, et al.
Publicado: (2023)
por: Choi, Haram, et al.
Publicado: (2023)
AdaIR: Exploiting Underlying Similarities of Image Restoration Tasks with Adapters
por: Chen, Hao-Wei, et al.
Publicado: (2024)
por: Chen, Hao-Wei, et al.
Publicado: (2024)
GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
por: Park, Seongheon, et al.
Publicado: (2025)
por: Park, Seongheon, et al.
Publicado: (2025)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
por: Hua, Zhenglin, et al.
Publicado: (2025)
por: Hua, Zhenglin, et al.
Publicado: (2025)
Finding the Correct Visual Evidence Without Forgetting: Mitigating Hallucination in LVLMs via Inter-Layer Visual Attention Discrepancy
por: Xie, Yutong, et al.
Publicado: (2026)
por: Xie, Yutong, et al.
Publicado: (2026)
Upcycling Text-to-Image Diffusion Models for Multi-Task Capabilities
por: Chavhan, Ruchika, et al.
Publicado: (2025)
por: Chavhan, Ruchika, et al.
Publicado: (2025)
Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks
por: Li, Chenjun
Publicado: (2026)
por: Li, Chenjun
Publicado: (2026)
Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding
por: Kwon, Mincheol, et al.
Publicado: (2026)
por: Kwon, Mincheol, et al.
Publicado: (2026)
Contribution-based Low-Rank Adaptation with Pre-training Model for Real Image Restoration
por: Park, Donwon, et al.
Publicado: (2024)
por: Park, Donwon, et al.
Publicado: (2024)
SFUOD: Source-Free Unknown Object Detection
por: Park, Keon-Hee, et al.
Publicado: (2025)
por: Park, Keon-Hee, et al.
Publicado: (2025)
Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)
por: Manevich, Avshalom, et al.
Publicado: (2024)
por: Manevich, Avshalom, et al.
Publicado: (2024)
Anomaly Detection by Effectively Leveraging Synthetic Images
por: Kang, Sungho, et al.
Publicado: (2025)
por: Kang, Sungho, et al.
Publicado: (2025)
Ejemplares similares
-
ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
por: Park, Yeji, et al.
Publicado: (2024) -
Enhancing Multi-Image Understanding through Delimiter Token Scaling
por: Lee, Minyoung, et al.
Publicado: (2026) -
LMLT: Low-to-high Multi-Level Vision Transformer for Image Super-Resolution
por: Kim, Jeongsoo, et al.
Publicado: (2024) -
Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs
por: Lee, Insu, et al.
Publicado: (2025) -
Sampling Bag of Views for Open-Vocabulary Object Detection
por: Choi, Hojun, et al.
Publicado: (2024)