Beyond Intermediate States: Explaining Visual Redundancy through Language
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Dingchen, Cao, Bowen, Zhang, Anran, Gu, Weibo, Hu, Winston, Chen, Guang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pensieve: Retrospect-then-Compare Mitigates Visual Hallucination
by: Yang, Dingchen, et al.
Published: (2024)
by: Yang, Dingchen, et al.
Published: (2024)
Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction
by: Jian, Yiren, et al.
Published: (2023)
by: Jian, Yiren, et al.
Published: (2023)
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers
by: Zhang, Renshan, et al.
Published: (2025)
by: Zhang, Renshan, et al.
Published: (2025)
A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models
by: Li, Duo, et al.
Published: (2025)
by: Li, Duo, et al.
Published: (2025)
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
by: Li, Hongliang, et al.
Published: (2025)
by: Li, Hongliang, et al.
Published: (2025)
Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models
by: Zhou, Haoran, et al.
Published: (2025)
by: Zhou, Haoran, et al.
Published: (2025)
Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
by: Dong, Yuhao, et al.
Published: (2024)
by: Dong, Yuhao, et al.
Published: (2024)
Color-Oriented Redundancy Reduction in Dataset Distillation
by: Yuan, Bowen, et al.
Published: (2024)
by: Yuan, Bowen, et al.
Published: (2024)
PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
by: Xing, Long, et al.
Published: (2024)
by: Xing, Long, et al.
Published: (2024)
Transfer Attack for Bad and Good: Explain and Boost Adversarial Transferability across Multimodal Large Language Models
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations
by: Wu, Xuecheng, et al.
Published: (2025)
by: Wu, Xuecheng, et al.
Published: (2025)
dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
by: Li, Yumeng, et al.
Published: (2025)
by: Li, Yumeng, et al.
Published: (2025)
MFil-Mamba: Multi-Filter Scanning for Spatial Redundancy-Aware Visual State Space Models
by: Khadka, Puskal, et al.
Published: (2026)
by: Khadka, Puskal, et al.
Published: (2026)
LWGANet: Addressing Spatial and Channel Redundancy in Remote Sensing Visual Tasks with Light-Weight Grouped Attention
by: Lu, Wei, et al.
Published: (2025)
by: Lu, Wei, et al.
Published: (2025)
Beyond Visual Cues: Synchronously Exploring Target-Centric Semantics for Vision-Language Tracking
by: Ge, Jiawei, et al.
Published: (2023)
by: Ge, Jiawei, et al.
Published: (2023)
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy
by: Yang, Te, et al.
Published: (2024)
by: Yang, Te, et al.
Published: (2024)
Efficient Redundancy Reduction for Open-Vocabulary Semantic Segmentation
by: Chen, Lin, et al.
Published: (2025)
by: Chen, Lin, et al.
Published: (2025)
AdaFV: Rethinking of Visual-Language alignment for VLM acceleration
by: Han, Jiayi, et al.
Published: (2025)
by: Han, Jiayi, et al.
Published: (2025)
Beyond the Visible: Disocclusion-Aware Editing via Proxy Dynamic Graphs
by: Qi, Anran, et al.
Published: (2025)
by: Qi, Anran, et al.
Published: (2025)
Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning
by: Pang, Yuqi, et al.
Published: (2025)
by: Pang, Yuqi, et al.
Published: (2025)
BabyVision: Visual Reasoning Beyond Language
by: Chen, Liang, et al.
Published: (2026)
by: Chen, Liang, et al.
Published: (2026)
Taming the Entropy Cliff: Variable Codebook Size Quantization for Autoregressive Visual Generation
by: Zheng, Bowen, et al.
Published: (2026)
by: Zheng, Bowen, et al.
Published: (2026)
Rolling Shutter Correction with Intermediate Distortion Flow Estimation
by: Cao, Mingdeng, et al.
Published: (2024)
by: Cao, Mingdeng, et al.
Published: (2024)
Video-Language Alignment via Spatio-Temporal Graph Transformer
by: Zhang, Shi-Xue, et al.
Published: (2024)
by: Zhang, Shi-Xue, et al.
Published: (2024)
VIHD: Visual Intervention-based Hallucination Detection for Medical Visual Question Answering
by: Chen, Jiayi, et al.
Published: (2026)
by: Chen, Jiayi, et al.
Published: (2026)
EmoSEM: Segment and Explain Emotion Stimuli in Visual Art
by: Zhang, Jing, et al.
Published: (2025)
by: Zhang, Jing, et al.
Published: (2025)
Beyond Degradation Redundancy: Contrastive Prompt Learning for All-in-One Image Restoration
by: Wu, Gang, et al.
Published: (2025)
by: Wu, Gang, et al.
Published: (2025)
Seeing Beyond Redundancy: Task Complexity's Role in Vision Token Specialization in VLLMs
by: Hannan, Darryl, et al.
Published: (2026)
by: Hannan, Darryl, et al.
Published: (2026)
TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
by: Yao, Linli, et al.
Published: (2025)
by: Yao, Linli, et al.
Published: (2025)
Discovering Intrinsic Spatial-Temporal Logic Rules to Explain Human Actions
by: Cao, Chengzhi, et al.
Published: (2023)
by: Cao, Chengzhi, et al.
Published: (2023)
FontCLIP: A Semantic Typography Visual-Language Model for Multilingual Font Applications
by: Tatsukawa, Yuki, et al.
Published: (2024)
by: Tatsukawa, Yuki, et al.
Published: (2024)
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
by: Gao, Mingjian, et al.
Published: (2026)
by: Gao, Mingjian, et al.
Published: (2026)
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
by: Wang, Yizhou, et al.
Published: (2025)
by: Wang, Yizhou, et al.
Published: (2025)
Enhancing Targeted Adversarial Attacks on Large Vision-Language Models via Intermediate Projector
by: Cao, Yiming, et al.
Published: (2025)
by: Cao, Yiming, et al.
Published: (2025)
GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval
by: Yang, Bowen, et al.
Published: (2025)
by: Yang, Bowen, et al.
Published: (2025)
Depicting Beyond Scores: Advancing Image Quality Assessment through Multi-modal Language Models
by: You, Zhiyuan, et al.
Published: (2023)
by: You, Zhiyuan, et al.
Published: (2023)
Explaining Representation by Mutual Information
by: Gu, Lifeng
Published: (2021)
by: Gu, Lifeng
Published: (2021)
VAEVQ: Enhancing Discrete Visual Tokenization through Variational Modeling
by: Yang, Sicheng, et al.
Published: (2025)
by: Yang, Sicheng, et al.
Published: (2025)
Language Models Can Explain Visual Features via Steering
by: Ferrando, Javier, et al.
Published: (2026)
by: Ferrando, Javier, et al.
Published: (2026)
Learning to Rank Patches for Unbiased Image Redundancy Reduction
by: Luo, Yang, et al.
Published: (2024)
by: Luo, Yang, et al.
Published: (2024)
Similar Items
-
Pensieve: Retrospect-then-Compare Mitigates Visual Hallucination
by: Yang, Dingchen, et al.
Published: (2024) -
Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction
by: Jian, Yiren, et al.
Published: (2023) -
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers
by: Zhang, Renshan, et al.
Published: (2025) -
A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models
by: Li, Duo, et al.
Published: (2025) -
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
by: Li, Hongliang, et al.
Published: (2025)