The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zhuowei, Shi, Haizhou, Gao, Yunhe, Liu, Di, Wang, Zhenting, Chen, Yuxiao, Liu, Ting, Zhao, Long, Wang, Hao, Metaxas, Dimitris N. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering
von: Chatzoudis, Gerasimos, et al.
Veröffentlicht: (2025)
von: Chatzoudis, Gerasimos, et al.
Veröffentlicht: (2025)
Training Like a Medical Resident: Context-Prior Learning Toward Universal Medical Image Segmentation
von: Gao, Yunhe, et al.
Veröffentlicht: (2023)
von: Gao, Yunhe, et al.
Veröffentlicht: (2023)
Implicit In-context Learning
von: Li, Zhuowei, et al.
Veröffentlicht: (2024)
von: Li, Zhuowei, et al.
Veröffentlicht: (2024)
BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models
von: Wang, Yibin, et al.
Veröffentlicht: (2024)
von: Wang, Yibin, et al.
Veröffentlicht: (2024)
TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
von: Zhang, Tunyu, et al.
Veröffentlicht: (2025)
von: Zhang, Tunyu, et al.
Veröffentlicht: (2025)
Show and Segment: Universal Medical Image Segmentation via In-Context Learning
von: Gao, Yunhe, et al.
Veröffentlicht: (2025)
von: Gao, Yunhe, et al.
Veröffentlicht: (2025)
Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation
von: Gu, Difei, et al.
Veröffentlicht: (2025)
von: Gu, Difei, et al.
Veröffentlicht: (2025)
Aligning Human Knowledge with Visual Concepts Towards Explainable Medical Image Classification
von: Gao, Yunhe, et al.
Veröffentlicht: (2024)
von: Gao, Yunhe, et al.
Veröffentlicht: (2024)
Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction
von: Zhao, Shiyu, et al.
Veröffentlicht: (2024)
von: Zhao, Shiyu, et al.
Veröffentlicht: (2024)
LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment
von: Gu, Difei, et al.
Veröffentlicht: (2025)
von: Gu, Difei, et al.
Veröffentlicht: (2025)
Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models
von: Dafnis, Konstantinos M., et al.
Veröffentlicht: (2025)
von: Dafnis, Konstantinos M., et al.
Veröffentlicht: (2025)
Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
von: Chatzoudis, Gerasimos, et al.
Veröffentlicht: (2026)
von: Chatzoudis, Gerasimos, et al.
Veröffentlicht: (2026)
Reducing Hallucinations in Vision-Language Models via Latent Space Steering
von: Liu, Sheng, et al.
Veröffentlicht: (2024)
von: Liu, Sheng, et al.
Veröffentlicht: (2024)
LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation
von: Jin, Can, et al.
Veröffentlicht: (2025)
von: Jin, Can, et al.
Veröffentlicht: (2025)
DIAGNOSIS: Detecting Unauthorized Data Usages in Text-to-image Diffusion Models
von: Wang, Zhenting, et al.
Veröffentlicht: (2023)
von: Wang, Zhenting, et al.
Veröffentlicht: (2023)
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
von: Li, Qiming, et al.
Veröffentlicht: (2026)
von: Li, Qiming, et al.
Veröffentlicht: (2026)
Token-Controlled Re-ranking for Sequential Recommendation via LLMs
von: Dai, Wenxi, et al.
Veröffentlicht: (2025)
von: Dai, Wenxi, et al.
Veröffentlicht: (2025)
Few-Step Diffusion Language Models via Trajectory Self-Distillation
von: Zhang, Tunyu, et al.
Veröffentlicht: (2026)
von: Zhang, Tunyu, et al.
Veröffentlicht: (2026)
How to Trace Latent Generative Model Generated Images without Artificial Watermark?
von: Wang, Zhenting, et al.
Veröffentlicht: (2024)
von: Wang, Zhenting, et al.
Veröffentlicht: (2024)
Steering Rectified Flow Models in the Vector Field for Controlled Image Generation
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
VerSe: Integrating Multiple Queries as Prompts for Versatile Cardiac MRI Segmentation
von: Guo, Bangwei, et al.
Veröffentlicht: (2024)
von: Guo, Bangwei, et al.
Veröffentlicht: (2024)
Beyond Explicit Edges: Robust Reasoning over Noisy and Sparse Knowledge Graphs
von: Gao, Hang, et al.
Veröffentlicht: (2026)
von: Gao, Hang, et al.
Veröffentlicht: (2026)
Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
von: Zhang, Yuanhong, et al.
Veröffentlicht: (2026)
von: Zhang, Yuanhong, et al.
Veröffentlicht: (2026)
Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models
von: Wu, Jialin, et al.
Veröffentlicht: (2026)
von: Wu, Jialin, et al.
Veröffentlicht: (2026)
On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models
von: Seo, Hoigi, et al.
Veröffentlicht: (2025)
von: Seo, Hoigi, et al.
Veröffentlicht: (2025)
LUCID-SAE: Learning Unified Vision-Language Sparse Codes for Interpretable Concept Discovery
von: Gu, Difei, et al.
Veröffentlicht: (2026)
von: Gu, Difei, et al.
Veröffentlicht: (2026)
Instantaneous Perception of Moving Objects in 3D
von: Liu, Di, et al.
Veröffentlicht: (2024)
von: Liu, Di, et al.
Veröffentlicht: (2024)
MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
von: Dao, Quan, et al.
Veröffentlicht: (2026)
von: Dao, Quan, et al.
Veröffentlicht: (2026)
On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression
von: Zhang, Xinwei, et al.
Veröffentlicht: (2026)
von: Zhang, Xinwei, et al.
Veröffentlicht: (2026)
K-Prism: A Knowledge-Guided and Prompt Integrated Universal Medical Image Segmentation Model
von: Guo, Bangwei, et al.
Veröffentlicht: (2025)
von: Guo, Bangwei, et al.
Veröffentlicht: (2025)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
von: Wang, Zihu, et al.
Veröffentlicht: (2025)
von: Wang, Zihu, et al.
Veröffentlicht: (2025)
MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
Improving Visual Reasoning with Iterative Evidence Refinement
von: Shi, Zeru, et al.
Veröffentlicht: (2026)
von: Shi, Zeru, et al.
Veröffentlicht: (2026)
Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models
von: Zou, Zhengtao, et al.
Veröffentlicht: (2025)
von: Zou, Zhengtao, et al.
Veröffentlicht: (2025)
One-shot Optimized Steering Vector for Hallucination Mitigation for VLMs
von: Shi, Youxu, et al.
Veröffentlicht: (2026)
von: Shi, Youxu, et al.
Veröffentlicht: (2026)
Improved Training Technique for Latent Consistency Models
von: Dao, Quan, et al.
Veröffentlicht: (2025)
von: Dao, Quan, et al.
Veröffentlicht: (2025)
Resolving Inconsistent Semantics in Multi-Dataset Image Segmentation
von: Zhangli, Qilong, et al.
Veröffentlicht: (2024)
von: Zhangli, Qilong, et al.
Veröffentlicht: (2024)
APEER: Automatic Prompt Engineering Enhances Large Language Model Reranking
von: Jin, Can, et al.
Veröffentlicht: (2024)
von: Jin, Can, et al.
Veröffentlicht: (2024)
ConsistentRFT: Reducing Visual Hallucinations in Flow-based Reinforcement Fine-Tuning
von: Tan, Xiaofeng, et al.
Veröffentlicht: (2026)
von: Tan, Xiaofeng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering
von: Chatzoudis, Gerasimos, et al.
Veröffentlicht: (2025) -
Training Like a Medical Resident: Context-Prior Learning Toward Universal Medical Image Segmentation
von: Gao, Yunhe, et al.
Veröffentlicht: (2023) -
Implicit In-context Learning
von: Li, Zhuowei, et al.
Veröffentlicht: (2024) -
BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models
von: Wang, Yibin, et al.
Veröffentlicht: (2024) -
TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
von: Zhang, Tunyu, et al.
Veröffentlicht: (2025)