Rethinking Visual Token Reduction in LVLMs Under Cross-Modal Misalignment
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Rui, Wang, Yunke, Luo, Yong, Du, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unsupervised Cross-domain Pulmonary Nodule Detection without Source Data
by: Xu, Rui, et al.
Published: (2023)
by: Xu, Rui, et al.
Published: (2023)
MAEDiff: Masked Autoencoder-enhanced Diffusion Models for Unsupervised Anomaly Detection in Brain Images
by: Xu, Rui, et al.
Published: (2024)
by: Xu, Rui, et al.
Published: (2024)
EchoingPixels: Cross-Modal Adaptive Token Reduction for Efficient Audio-Visual LLMs
by: Gong, Chao, et al.
Published: (2025)
by: Gong, Chao, et al.
Published: (2025)
IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
by: Sun, Zhichao, et al.
Published: (2026)
by: Sun, Zhichao, et al.
Published: (2026)
MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality
by: Yang, Panqi, et al.
Published: (2026)
by: Yang, Panqi, et al.
Published: (2026)
Multi-Tailed Vision Transformer for Efficient Inference
by: Wang, Yunke, et al.
Published: (2022)
by: Wang, Yunke, et al.
Published: (2022)
Visual Imitation Learning with Calibrated Contrastive Representation
by: Wang, Yunke, et al.
Published: (2024)
by: Wang, Yunke, et al.
Published: (2024)
CollagePrompt: A Benchmark for Budget-Friendly Visual Recognition with GPT-4V
by: Xu, Siyu, et al.
Published: (2024)
by: Xu, Siyu, et al.
Published: (2024)
UniV2D: Bridging Visual Restoration and Semantic Perception for Underwater Salient Object Detection
by: Chang, Laibin, et al.
Published: (2026)
by: Chang, Laibin, et al.
Published: (2026)
FusionSAM: Visual Multi-Modal Learning with Segment Anything
by: Li, Daixun, et al.
Published: (2024)
by: Li, Daixun, et al.
Published: (2024)
VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
by: Zhang, Ce, et al.
Published: (2025)
by: Zhang, Ce, et al.
Published: (2025)
STMI: Segmentation-Guided Token Modulation with Cross-Modal Hypergraph Interaction for Multi-Modal Object Re-Identification
by: Xu, Xingguo, et al.
Published: (2026)
by: Xu, Xingguo, et al.
Published: (2026)
Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP
by: Herzog, Jonas, et al.
Published: (2026)
by: Herzog, Jonas, et al.
Published: (2026)
On the Value of Cross-Modal Misalignment in Multimodal Representation Learning
by: Cai, Yichao, et al.
Published: (2025)
by: Cai, Yichao, et al.
Published: (2025)
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination
by: Chen, Yangneng, et al.
Published: (2026)
by: Chen, Yangneng, et al.
Published: (2026)
SAVAA: Mitigating Hallucinations in LVLMs via Step-wise Adaptive Visual Attention Amplification
by: Zhang, Jiacheng, et al.
Published: (2026)
by: Zhang, Jiacheng, et al.
Published: (2026)
Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing
by: He, Jingxuan, et al.
Published: (2026)
by: He, Jingxuan, et al.
Published: (2026)
Rethinking Token Reduction for Large Vision-Language Models
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
CORE: Compact Object-centric REpresentations as a New Paradigm for Token Merging in LVLMs
by: Lei, Jingyu, et al.
Published: (2025)
by: Lei, Jingyu, et al.
Published: (2025)
Evidence Packing for Cross-Domain Image Deepfake Detection with LVLMs
by: Liu, Yuxin, et al.
Published: (2026)
by: Liu, Yuxin, et al.
Published: (2026)
Rethinking Token Reduction for Diffusion Models via Output-Similarity-Awareness
by: Lee, Hangyeol, et al.
Published: (2026)
by: Lee, Hangyeol, et al.
Published: (2026)
MambaScope: Coarse-to-Fine Scoping for Efficient Vision Mamba
by: Liu, Shanhui, et al.
Published: (2025)
by: Liu, Shanhui, et al.
Published: (2025)
Self-Prophetic Decoding to Unlock Visual Search in LVLMs
by: He, Zhendong, et al.
Published: (2026)
by: He, Zhendong, et al.
Published: (2026)
Robust Self-Supervised Cross-Modal Super-Resolution against Real-World Misaligned Observations
by: Dong, Xiaoyu, et al.
Published: (2026)
by: Dong, Xiaoyu, et al.
Published: (2026)
Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
by: Lv, Zhengyao, et al.
Published: (2025)
by: Lv, Zhengyao, et al.
Published: (2025)
CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs
by: Kan, Zhehan, et al.
Published: (2024)
by: Kan, Zhehan, et al.
Published: (2024)
Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal Learning
by: Wang, Alex Jinpeng, et al.
Published: (2024)
by: Wang, Alex Jinpeng, et al.
Published: (2024)
Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
by: Hu, Qingguo, et al.
Published: (2025)
by: Hu, Qingguo, et al.
Published: (2025)
Marine Saliency Segmenter: Object-Focused Conditional Diffusion with Region-Level Semantic Knowledge Distillation
by: Chang, Laibin, et al.
Published: (2025)
by: Chang, Laibin, et al.
Published: (2025)
Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives
by: Zhang, Zeliang, et al.
Published: (2025)
by: Zhang, Zeliang, et al.
Published: (2025)
Visible-Infrared Person Re-Identification via Patch-Mixed Cross-Modality Learning
by: Qian, Zhihao, et al.
Published: (2023)
by: Qian, Zhihao, et al.
Published: (2023)
REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
by: He, Qiyuan, et al.
Published: (2025)
by: He, Qiyuan, et al.
Published: (2025)
Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement
by: Qin, Zhenxin, et al.
Published: (2026)
by: Qin, Zhenxin, et al.
Published: (2026)
MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs
by: Chen, Huiyi, et al.
Published: (2025)
by: Chen, Huiyi, et al.
Published: (2025)
HERO: Rethinking Visual Token Early Dropping in High-Resolution Large Vision-Language Models
by: Li, Xu, et al.
Published: (2025)
by: Li, Xu, et al.
Published: (2025)
See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
by: Dai, Ziyun, et al.
Published: (2025)
by: Dai, Ziyun, et al.
Published: (2025)
AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
by: Zhang, Yuan, et al.
Published: (2025)
by: Zhang, Yuan, et al.
Published: (2025)
Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance
by: Shen, Weijie, et al.
Published: (2025)
by: Shen, Weijie, et al.
Published: (2025)
TinyChemVL: Advancing Chemical Vision-Language Models via Efficient Visual Token Reduction and Complex Reaction Tasks
by: Zhao, Xuanle, et al.
Published: (2025)
by: Zhao, Xuanle, et al.
Published: (2025)
Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens
by: Zheng, Haohan, et al.
Published: (2025)
by: Zheng, Haohan, et al.
Published: (2025)
Similar Items
-
Unsupervised Cross-domain Pulmonary Nodule Detection without Source Data
by: Xu, Rui, et al.
Published: (2023) -
MAEDiff: Masked Autoencoder-enhanced Diffusion Models for Unsupervised Anomaly Detection in Brain Images
by: Xu, Rui, et al.
Published: (2024) -
EchoingPixels: Cross-Modal Adaptive Token Reduction for Efficient Audio-Visual LLMs
by: Gong, Chao, et al.
Published: (2025) -
IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
by: Sun, Zhichao, et al.
Published: (2026) -
MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality
by: Yang, Panqi, et al.
Published: (2026)