MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Chao, Yang, Jianming, Zhou, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Conscious Gaze: Adaptive Attention Mechanisms for Hallucination Mitigation in Vision-Language Models
von: Bu, Weijue, et al.
Veröffentlicht: (2025)
von: Bu, Weijue, et al.
Veröffentlicht: (2025)
Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens
von: Shen, Meng, et al.
Veröffentlicht: (2026)
von: Shen, Meng, et al.
Veröffentlicht: (2026)
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
von: Lee, Kyuho, et al.
Veröffentlicht: (2025)
von: Lee, Kyuho, et al.
Veröffentlicht: (2025)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
BlindSight: Harnessing Sparsity for Efficient Vision-Language Models
von: Srikrishnan, Tharun Adithya, et al.
Veröffentlicht: (2025)
von: Srikrishnan, Tharun Adithya, et al.
Veröffentlicht: (2025)
Textual and Visual Guided Task Adaptation for Source-Free Cross-Domain Few-Shot Segmentation
von: Liu, Jianming, et al.
Veröffentlicht: (2025)
von: Liu, Jianming, et al.
Veröffentlicht: (2025)
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
von: Zeng, Zhitao, et al.
Veröffentlicht: (2025)
von: Zeng, Zhitao, et al.
Veröffentlicht: (2025)
A Vision-Language Model for Focal Liver Lesion Classification
von: Jian, Song, et al.
Veröffentlicht: (2025)
von: Jian, Song, et al.
Veröffentlicht: (2025)
Towards Hard and Soft Shadow Removal via Dual-Branch Separation Network and Vision Transformer
von: Liang, Jiajia
Veröffentlicht: (2025)
von: Liang, Jiajia
Veröffentlicht: (2025)
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
von: Feng, Yigui, et al.
Veröffentlicht: (2026)
von: Feng, Yigui, et al.
Veröffentlicht: (2026)
More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage
von: He, Wei
Veröffentlicht: (2026)
von: He, Wei
Veröffentlicht: (2026)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
von: Li, Huibin, et al.
Veröffentlicht: (2025)
von: Li, Huibin, et al.
Veröffentlicht: (2025)
NAC-TCN: Temporal Convolutional Networks with Causal Dilated Neighborhood Attention for Emotion Understanding
von: Mehta, Alexander, et al.
Veröffentlicht: (2023)
von: Mehta, Alexander, et al.
Veröffentlicht: (2023)
OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance
von: Wang, Chaoyi, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyi, et al.
Veröffentlicht: (2025)
Context-Aware Indoor Point Cloud Object Generation through User Instructions
von: Luo, Yiyang, et al.
Veröffentlicht: (2023)
von: Luo, Yiyang, et al.
Veröffentlicht: (2023)
Computer Vision for Clinical Gait Analysis: A Gait Abnormality Video Dataset
von: Ranjan, Rahm, et al.
Veröffentlicht: (2024)
von: Ranjan, Rahm, et al.
Veröffentlicht: (2024)
ViG-LRGC: Vision Graph Neural Networks with Learnable Reparameterized Graph Construction
von: Elsharkawi, Ismael, et al.
Veröffentlicht: (2025)
von: Elsharkawi, Ismael, et al.
Veröffentlicht: (2025)
Mitigating Catastrophic Forgetting in the Incremental Learning of Medical Images
von: Yavari, Sara, et al.
Veröffentlicht: (2025)
von: Yavari, Sara, et al.
Veröffentlicht: (2025)
Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models
von: Dubois, L'ea, et al.
Veröffentlicht: (2025)
von: Dubois, L'ea, et al.
Veröffentlicht: (2025)
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
von: Han, Yudong, et al.
Veröffentlicht: (2026)
von: Han, Yudong, et al.
Veröffentlicht: (2026)
Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models
von: Freitas, Diogo, et al.
Veröffentlicht: (2025)
von: Freitas, Diogo, et al.
Veröffentlicht: (2025)
A Comprehensive Review of YOLO Architectures in Computer Vision: From YOLOv1 to YOLOv8 and YOLO-NAS
von: Terven, Juan, et al.
Veröffentlicht: (2023)
von: Terven, Juan, et al.
Veröffentlicht: (2023)
One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual Representation
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
von: Rudman, William, et al.
Veröffentlicht: (2026)
von: Rudman, William, et al.
Veröffentlicht: (2026)
Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
von: Ji, Yikun, et al.
Veröffentlicht: (2025)
von: Ji, Yikun, et al.
Veröffentlicht: (2025)
FocusedAD: Character-centric Movie Audio Description
von: Ye, Xiaojun, et al.
Veröffentlicht: (2025)
von: Ye, Xiaojun, et al.
Veröffentlicht: (2025)
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
von: Mehta, Vinit, et al.
Veröffentlicht: (2025)
von: Mehta, Vinit, et al.
Veröffentlicht: (2025)
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
von: Han, Yudong, et al.
Veröffentlicht: (2024)
von: Han, Yudong, et al.
Veröffentlicht: (2024)
Revisiting SVD and Wavelet Difference Reduction for Lossy Image Compression: A Reproducibility Study
von: Makarova, Alena
Veröffentlicht: (2025)
von: Makarova, Alena
Veröffentlicht: (2025)
3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding
von: Chen, Yiping, et al.
Veröffentlicht: (2026)
von: Chen, Yiping, et al.
Veröffentlicht: (2026)
Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos
von: Li, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyang, et al.
Veröffentlicht: (2025)
A Challenging Benchmark of Anime Style Recognition
von: Li, Haotang, et al.
Veröffentlicht: (2022)
von: Li, Haotang, et al.
Veröffentlicht: (2022)
Breaking the Resource Wall: Geometry-Guided Sequence Modeling for Efficient Semantic Segmentation
von: Chan, Sheng-Wei, et al.
Veröffentlicht: (2026)
von: Chan, Sheng-Wei, et al.
Veröffentlicht: (2026)
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
RGB-Only Gaussian Splatting SLAM for Unbounded Outdoor Scenes
von: Yu, Sicheng, et al.
Veröffentlicht: (2025)
von: Yu, Sicheng, et al.
Veröffentlicht: (2025)
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
von: Wang, Zhaohui, et al.
Veröffentlicht: (2025)
von: Wang, Zhaohui, et al.
Veröffentlicht: (2025)
ERNet: Efficient Non-Rigid Registration Network for Point Sequences
von: He, Guangzhao, et al.
Veröffentlicht: (2025)
von: He, Guangzhao, et al.
Veröffentlicht: (2025)
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
von: Su, Yuetong, et al.
Veröffentlicht: (2025)
von: Su, Yuetong, et al.
Veröffentlicht: (2025)
LEGO: LoRA-Enabled Generator-Oriented Framework for Synthetic Image Detection
von: Xiao, Yutong, et al.
Veröffentlicht: (2026)
von: Xiao, Yutong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Conscious Gaze: Adaptive Attention Mechanisms for Hallucination Mitigation in Vision-Language Models
von: Bu, Weijue, et al.
Veröffentlicht: (2025) -
Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens
von: Shen, Meng, et al.
Veröffentlicht: (2026) -
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
von: Lee, Kyuho, et al.
Veröffentlicht: (2025) -
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2024) -
BlindSight: Harnessing Sparsity for Efficient Vision-Language Models
von: Srikrishnan, Tharun Adithya, et al.
Veröffentlicht: (2025)