Stop learning it all to mitigate visual hallucination, Focus on the hallucination target
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yoon, Dokyoon, Song, Youngsook, Park, Woomyong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset
von: Liu, Jiazhen, et al.
Veröffentlicht: (2024)
von: Liu, Jiazhen, et al.
Veröffentlicht: (2024)
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
von: Zhang, Xinsong, et al.
Veröffentlicht: (2025)
von: Zhang, Xinsong, et al.
Veröffentlicht: (2025)
When normalization hallucinates: unseen risks in AI-powered whole slide image processing
von: Moens, Karel, et al.
Veröffentlicht: (2025)
von: Moens, Karel, et al.
Veröffentlicht: (2025)
RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
Attention to details, logits to truth: visual-aware attention and logits enhancement to mitigate hallucinations in LVLMs
von: Wang, Jingyi, et al.
Veröffentlicht: (2026)
von: Wang, Jingyi, et al.
Veröffentlicht: (2026)
VCE: A zero-cost hallucination mitigation method of LVLMs via visual contrastive editing
von: Huang, Yanbin, et al.
Veröffentlicht: (2026)
von: Huang, Yanbin, et al.
Veröffentlicht: (2026)
Better Safe Than Sorry? Overreaction Problem of Vision Language Models in Visual Emergency Recognition
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
sFRC for assessing hallucinations in medical image restoration
von: Kc, Prabhat, et al.
Veröffentlicht: (2026)
von: Kc, Prabhat, et al.
Veröffentlicht: (2026)
Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding
von: Kwon, Mincheol, et al.
Veröffentlicht: (2026)
von: Kwon, Mincheol, et al.
Veröffentlicht: (2026)
Focus Matters: Phase-Aware Suppression for Hallucination in Vision-Language Models
von: Kim, Sohyeon, et al.
Veröffentlicht: (2026)
von: Kim, Sohyeon, et al.
Veröffentlicht: (2026)
Visual hallucination detection in large vision-language models via evidential conflict
von: Huang, Tao, et al.
Veröffentlicht: (2025)
von: Huang, Tao, et al.
Veröffentlicht: (2025)
Extend3D: Town-Scale 3D Generation
von: Yoon, Seungwoo, et al.
Veröffentlicht: (2026)
von: Yoon, Seungwoo, et al.
Veröffentlicht: (2026)
Self-supervised visual learning for analyzing firearms trafficking activities on the Web
von: Konstantakos, Sotirios, et al.
Veröffentlicht: (2023)
von: Konstantakos, Sotirios, et al.
Veröffentlicht: (2023)
Your One-Stop Solution for AI-Generated Video Detection
von: Ma, Long, et al.
Veröffentlicht: (2026)
von: Ma, Long, et al.
Veröffentlicht: (2026)
Multi-agent Long-term 3D Human Pose Forecasting via Interaction-aware Trajectory Conditioning
von: Jeong, Jaewoo, et al.
Veröffentlicht: (2024)
von: Jeong, Jaewoo, et al.
Veröffentlicht: (2024)
Focus on Focus: Focus-oriented Representation Learning and Multi-view Cross-modal Alignment for Glioma Grading
von: Pan, Li, et al.
Veröffentlicht: (2024)
von: Pan, Li, et al.
Veröffentlicht: (2024)
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models
von: Choi, Dasol, et al.
Veröffentlicht: (2026)
von: Choi, Dasol, et al.
Veröffentlicht: (2026)
Point-RTD: Replaced Token Denoising for Pretraining Transformer Models on Point Clouds
von: Stone, Gunner, et al.
Veröffentlicht: (2025)
von: Stone, Gunner, et al.
Veröffentlicht: (2025)
GaussianFocus: Constrained Attention Focus for 3D Gaussian Splatting
von: Huang, Zexu, et al.
Veröffentlicht: (2025)
von: Huang, Zexu, et al.
Veröffentlicht: (2025)
SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination
von: Park, Sangha, et al.
Veröffentlicht: (2025)
von: Park, Sangha, et al.
Veröffentlicht: (2025)
Semantic Diversity-aware Prototype-based Learning for Unbiased Scene Graph Generation
von: Jeon, Jaehyeong, et al.
Veröffentlicht: (2024)
von: Jeon, Jaehyeong, et al.
Veröffentlicht: (2024)
CMI-MTL: Cross-Mamba interaction based multi-task learning for medical visual question answering
von: Jin, Qiangguo, et al.
Veröffentlicht: (2025)
von: Jin, Qiangguo, et al.
Veröffentlicht: (2025)
Triggering hallucinations in model-based MRI reconstruction via adversarial perturbations
von: Buğday, Suna, et al.
Veröffentlicht: (2026)
von: Buğday, Suna, et al.
Veröffentlicht: (2026)
REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training
von: Wang, Ziqiao, et al.
Veröffentlicht: (2025)
von: Wang, Ziqiao, et al.
Veröffentlicht: (2025)
Focusable Monocular Depth Estimation
von: Du, Yuxin, et al.
Veröffentlicht: (2026)
von: Du, Yuxin, et al.
Veröffentlicht: (2026)
Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity
von: Jung, Jaeyoon, et al.
Veröffentlicht: (2026)
von: Jung, Jaeyoon, et al.
Veröffentlicht: (2026)
Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
Unsupervised Homography Estimation on Multimodal Image Pair via Alternating Optimization
von: Song, Sanghyeob, et al.
Veröffentlicht: (2024)
von: Song, Sanghyeob, et al.
Veröffentlicht: (2024)
JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images
von: Wang, Zhecan, et al.
Veröffentlicht: (2024)
von: Wang, Zhecan, et al.
Veröffentlicht: (2024)
FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus
von: Jin, Qiaoqiao, et al.
Veröffentlicht: (2025)
von: Jin, Qiaoqiao, et al.
Veröffentlicht: (2025)
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
von: Park, Sangha, et al.
Veröffentlicht: (2025)
von: Park, Sangha, et al.
Veröffentlicht: (2025)
OV-MAP : Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots
von: Kim, Juno, et al.
Veröffentlicht: (2025)
von: Kim, Juno, et al.
Veröffentlicht: (2025)
MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding
von: Cao, Shiwen, et al.
Veröffentlicht: (2025)
von: Cao, Shiwen, et al.
Veröffentlicht: (2025)
Face Reconstruction Transfer Attack as Out-of-Distribution Generalization
von: Jung, Yoon Gyo, et al.
Veröffentlicht: (2024)
von: Jung, Yoon Gyo, et al.
Veröffentlicht: (2024)
NextStop: An Improved Tracker For Panoptic LIDAR Segmentation Data
von: Alkalay, Nirit, et al.
Veröffentlicht: (2025)
von: Alkalay, Nirit, et al.
Veröffentlicht: (2025)
Cross-modality debiasing: using language to mitigate sub-population shifts in imaging
von: Pang, Yijiang, et al.
Veröffentlicht: (2024)
von: Pang, Yijiang, et al.
Veröffentlicht: (2024)
Adaptive Self-training Framework for Fine-grained Scene Graph Generation
von: Kim, Kibum, et al.
Veröffentlicht: (2024)
von: Kim, Kibum, et al.
Veröffentlicht: (2024)
Learning to Stop Overthinking at Test Time
von: Bao, Hieu Tran, et al.
Veröffentlicht: (2025)
von: Bao, Hieu Tran, et al.
Veröffentlicht: (2025)
Multi-Focused Video Group Activities Hashing
von: Qi, Zhongmiao, et al.
Veröffentlicht: (2025)
von: Qi, Zhongmiao, et al.
Veröffentlicht: (2025)
FENet: Focusing Enhanced Network for Lane Detection
von: Wang, Liman, et al.
Veröffentlicht: (2023)
von: Wang, Liman, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset
von: Liu, Jiazhen, et al.
Veröffentlicht: (2024) -
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
von: Zhang, Xinsong, et al.
Veröffentlicht: (2025) -
When normalization hallucinates: unseen risks in AI-powered whole slide image processing
von: Moens, Karel, et al.
Veröffentlicht: (2025) -
RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2024) -
Attention to details, logits to truth: visual-aware attention and logits enhancement to mitigate hallucinations in LVLMs
von: Wang, Jingyi, et al.
Veröffentlicht: (2026)