Saved in:
| Main Authors: | Zhang, Xiaofeng, Zhu, Yuanchao, Gu, Chaochen, Yuan, Xiaosong, Zhao, Qiyan, Cao, Jiawei, Tang, Feilong, Fan, Sinan, Shen, Yaomin, Shen, Chen, Tang, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.20279 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models
by: Zhao, Qiyan, et al.
Published: (2025)
by: Zhao, Qiyan, et al.
Published: (2025)
Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
by: Zhang, Xiaofeng, et al.
Published: (2024)
by: Zhang, Xiaofeng, et al.
Published: (2024)
From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
by: Zhang, Xiaofeng, et al.
Published: (2024)
by: Zhang, Xiaofeng, et al.
Published: (2024)
Enlighten-Your-Voice: When Multimodal Meets Zero-shot Low-light Image Enhancement
by: Zhang, Xiaofeng, et al.
Published: (2023)
by: Zhang, Xiaofeng, et al.
Published: (2023)
A-MESS: Anchor based Multimodal Embedding with Semantic Synchronization for Multimodal Intent Recognition
by: Shen, Yaomin, et al.
Published: (2025)
by: Shen, Yaomin, et al.
Published: (2025)
Student-Oriented Teacher Knowledge Refinement for Knowledge Distillation
by: Shen, Chaomin, et al.
Published: (2024)
by: Shen, Chaomin, et al.
Published: (2024)
ConsensusDrop: Fusing Visual and Cross-Modal Saliency for Efficient Vision Language Models
by: Parikh, Dhruv, et al.
Published: (2026)
by: Parikh, Dhruv, et al.
Published: (2026)
Context Tokens are Anchors: Understanding the Repetition Curse in dMLLMs from an Information Flow Perspective
by: Zhao, Qiyan, et al.
Published: (2026)
by: Zhao, Qiyan, et al.
Published: (2026)
Harmonizing knowledge Transfer in Neural Network with Unified Distillation
by: Huang, Yaomin, et al.
Published: (2024)
by: Huang, Yaomin, et al.
Published: (2024)
NEARL-CLIP: Interacted Query Adaptation with Orthogonal Regularization for Medical Vision-Language Understanding
by: Peng, Zelin, et al.
Published: (2025)
by: Peng, Zelin, et al.
Published: (2025)
Saliency Guided Longitudinal Medical Visual Question Answering
by: Wu, Jialin, et al.
Published: (2025)
by: Wu, Jialin, et al.
Published: (2025)
Learning Where to Focus: Density-Driven Guidance for Detecting Dense Tiny Objects
by: Zhao, Zhicheng, et al.
Published: (2025)
by: Zhao, Zhicheng, et al.
Published: (2025)
SFC: Shared Feature Calibration in Weakly Supervised Semantic Segmentation
by: Zhao, Xinqiao, et al.
Published: (2024)
by: Zhao, Xinqiao, et al.
Published: (2024)
Exploring Saliency Bias in Manipulation Detection
by: Krinsky, Joshua, et al.
Published: (2024)
by: Krinsky, Joshua, et al.
Published: (2024)
Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
by: Li, Liyang, et al.
Published: (2026)
by: Li, Liyang, et al.
Published: (2026)
SASep: Saliency-Aware Structured Separation of Geometry and Feature for Open Set Learning on Point Clouds
by: Xu, Jinfeng, et al.
Published: (2025)
by: Xu, Jinfeng, et al.
Published: (2025)
SGP-SAM: Self-Gated Prompting for Transferring 3D Segment Anything Models to Lesion Segmentation
by: Tang, Zixuan, et al.
Published: (2026)
by: Tang, Zixuan, et al.
Published: (2026)
HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit
by: Wu, Hao, et al.
Published: (2026)
by: Wu, Hao, et al.
Published: (2026)
Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models
by: Qi, Jianing, et al.
Published: (2025)
by: Qi, Jianing, et al.
Published: (2025)
Who Can We Trust? Scope-Aware Video Moment Retrieval with Multi-Agent Conflict
by: Wu, Chaochen, et al.
Published: (2025)
by: Wu, Chaochen, et al.
Published: (2025)
Learning Compact Vision Tokens for Efficient Large Multimodal Models
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Robust Mesh Saliency Ground Truth Acquisition in VR via View Cone Sampling and Manifold Diffusion
by: Zheng, Guoquan, et al.
Published: (2026)
by: Zheng, Guoquan, et al.
Published: (2026)
Fine-grained Image Retrieval via Dual-Vision Adaptation
by: Jiang, Xin, et al.
Published: (2025)
by: Jiang, Xin, et al.
Published: (2025)
DISSECT: Diagnosing Where Vision Ends and Language Priors Begin in Scientific VLMs
by: Kukreja, Dikshant, et al.
Published: (2026)
by: Kukreja, Dikshant, et al.
Published: (2026)
RobustMerge: Parameter-Efficient Model Merging for MLLMs with Direction Robustness
by: Zeng, Fanhu, et al.
Published: (2025)
by: Zeng, Fanhu, et al.
Published: (2025)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
Bridge then Begin Anew: Generating Target-relevant Intermediate Model for Source-free Visual Emotion Adaptation
by: Zhu, Jiankun, et al.
Published: (2024)
by: Zhu, Jiankun, et al.
Published: (2024)
Instance-Warp: Saliency Guided Image Warping for Unsupervised Domain Adaptation
by: Zheng, Shen, et al.
Published: (2024)
by: Zheng, Shen, et al.
Published: (2024)
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
by: Song, Jiale, et al.
Published: (2026)
by: Song, Jiale, et al.
Published: (2026)
Advancing Weight and Channel Sparsification with Enhanced Saliency
by: Sun, Xinglong, et al.
Published: (2025)
by: Sun, Xinglong, et al.
Published: (2025)
Focus on Low-Resolution Information: Multi-Granular Information-Lossless Model for Low-Resolution Human Pose Estimation
by: Gu, Zejun, et al.
Published: (2024)
by: Gu, Zejun, et al.
Published: (2024)
TextMamba: Scene Text Detector with Mamba
by: Zhao, Qiyan, et al.
Published: (2025)
by: Zhao, Qiyan, et al.
Published: (2025)
EvRainDrop: HyperGraph-guided Completion for Effective Frame and Event Stream Aggregation
by: Wang, Futian, et al.
Published: (2025)
by: Wang, Futian, et al.
Published: (2025)
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
by: Tang, Feilong, et al.
Published: (2025)
by: Tang, Feilong, et al.
Published: (2025)
C^2ROPE: Causal Continuous Rotary Positional Encoding for 3D Large Multimodal-Models Reasoning
by: Ye, Guanting, et al.
Published: (2026)
by: Ye, Guanting, et al.
Published: (2026)
SSiT: Saliency-guided Self-supervised Image Transformer for Diabetic Retinopathy Grading
by: Huang, Yijin, et al.
Published: (2022)
by: Huang, Yijin, et al.
Published: (2022)
Saliency-Bench: A Comprehensive Benchmark for Evaluating Visual Explanations
by: Zhang, Yifei, et al.
Published: (2023)
by: Zhang, Yifei, et al.
Published: (2023)
PairDropGS: Paired Dropout-Induced Consistency Regularization for Sparse-View Gaussian Splatting
by: Li, Hantang, et al.
Published: (2026)
by: Li, Hantang, et al.
Published: (2026)
Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization
by: Lyu, Xinyu, et al.
Published: (2024)
by: Lyu, Xinyu, et al.
Published: (2024)
Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding
by: Xu, Zhongxing, et al.
Published: (2026)
by: Xu, Zhongxing, et al.
Published: (2026)
Similar Items
-
MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models
by: Zhao, Qiyan, et al.
Published: (2025) -
Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
by: Zhang, Xiaofeng, et al.
Published: (2024) -
From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
by: Zhang, Xiaofeng, et al.
Published: (2024) -
Enlighten-Your-Voice: When Multimodal Meets Zero-shot Low-light Image Enhancement
by: Zhang, Xiaofeng, et al.
Published: (2023) -
A-MESS: Anchor based Multimodal Embedding with Semantic Synchronization for Multimodal Intent Recognition
by: Shen, Yaomin, et al.
Published: (2025)