Saved in:
| Main Authors: | Wang, Yu, Li, Sharon |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.13403 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Representation Forcing for Bottleneck-Free Unified Multimodal Models
by: Wang, Yuqing, et al.
Published: (2026)
by: Wang, Yuqing, et al.
Published: (2026)
Information Bottleneck Approach to Spatial Attention Learning
by: Lai, Qiuxia, et al.
Published: (2021)
by: Lai, Qiuxia, et al.
Published: (2021)
Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers
by: Hong, Jung-Ho, et al.
Published: (2025)
by: Hong, Jung-Ho, et al.
Published: (2025)
Graph Integrated Multimodal Concept Bottleneck Model
by: Lin, Jiakai, et al.
Published: (2025)
by: Lin, Jiakai, et al.
Published: (2025)
Unveiling the Visual Counting Bottleneck in Vision-Language Models
by: Pang, Xingzhou, et al.
Published: (2026)
by: Pang, Xingzhou, et al.
Published: (2026)
Multimodal Conditional Information Bottleneck for Generalizable AI-Generated Image Detection
by: Qin, Haotian, et al.
Published: (2025)
by: Qin, Haotian, et al.
Published: (2025)
Breaking Dual Bottlenecks: Evolving Unified Multimodal Models into Self-Adaptive Interleaved Visual Reasoners
by: Liu, Qingyang, et al.
Published: (2026)
by: Liu, Qingyang, et al.
Published: (2026)
Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning
by: Liu, Shih-Wen, et al.
Published: (2025)
by: Liu, Shih-Wen, et al.
Published: (2025)
Prototypical Information Bottlenecking and Disentangling for Multimodal Cancer Survival Prediction
by: Zhang, Yilan, et al.
Published: (2024)
by: Zhang, Yilan, et al.
Published: (2024)
When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse
by: Huang, Yihuan, et al.
Published: (2026)
by: Huang, Yihuan, et al.
Published: (2026)
Peeking Behind the Curtains of Residual Learning
by: Zhang, Tunhou, et al.
Published: (2024)
by: Zhang, Tunhou, et al.
Published: (2024)
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge
by: Chen, Ruiming, et al.
Published: (2025)
by: Chen, Ruiming, et al.
Published: (2025)
HIFICL: High-Fidelity In-Context Learning for Multimodal Tasks
by: Li, Xiaoyu, et al.
Published: (2026)
by: Li, Xiaoyu, et al.
Published: (2026)
Generative Multimodal Models are In-Context Learners
by: Sun, Quan, et al.
Published: (2023)
by: Sun, Quan, et al.
Published: (2023)
Language Guided Concept Bottleneck Models for Interpretable Continual Learning
by: Yu, Lu, et al.
Published: (2025)
by: Yu, Lu, et al.
Published: (2025)
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
by: Endo, Mark, et al.
Published: (2025)
by: Endo, Mark, et al.
Published: (2025)
Towards Faithful Multimodal Concept Bottleneck Models
by: Moreau, Pierre, et al.
Published: (2026)
by: Moreau, Pierre, et al.
Published: (2026)
What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness
by: He, Yusheng, et al.
Published: (2026)
by: He, Yusheng, et al.
Published: (2026)
Towards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck Perspective
by: Long, Kaifang, et al.
Published: (2026)
by: Long, Kaifang, et al.
Published: (2026)
Learning Unsupervised Gaze Representation via Eye Mask Driven Information Bottleneck
by: Jiang, Yangzhou, et al.
Published: (2024)
by: Jiang, Yangzhou, et al.
Published: (2024)
Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
by: Li, Yanshu, et al.
Published: (2025)
by: Li, Yanshu, et al.
Published: (2025)
Partially Shared Concept Bottleneck Models
by: Zhao, Delong, et al.
Published: (2025)
by: Zhao, Delong, et al.
Published: (2025)
Personal Visual Context Learning in Large Multimodal Models
by: Xue, Zihui, et al.
Published: (2026)
by: Xue, Zihui, et al.
Published: (2026)
Towards Multimodal In-Context Learning for Vision & Language Models
by: Doveh, Sivan, et al.
Published: (2024)
by: Doveh, Sivan, et al.
Published: (2024)
Multi-scale Bottleneck Transformer for Weakly Supervised Multimodal Violence Detection
by: Sun, Shengyang, et al.
Published: (2024)
by: Sun, Shengyang, et al.
Published: (2024)
Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval?
by: Feng, Hengyi, et al.
Published: (2025)
by: Feng, Hengyi, et al.
Published: (2025)
Disentangled Representation Learning with Transmitted Information Bottleneck
by: Dang, Zhuohang, et al.
Published: (2023)
by: Dang, Zhuohang, et al.
Published: (2023)
CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
by: Li, Yanshu, et al.
Published: (2025)
by: Li, Yanshu, et al.
Published: (2025)
Can We Leave Deepfake Data Behind in Training Deepfake Detector?
by: Cheng, Jikang, et al.
Published: (2024)
by: Cheng, Jikang, et al.
Published: (2024)
MoVA: Adapting Mixture of Vision Experts to Multimodal Context
by: Zong, Zhuofan, et al.
Published: (2024)
by: Zong, Zhuofan, et al.
Published: (2024)
Information Bottleneck-Guided Heterogeneous Graph Learning for Interpretable Neurodevelopmental Disorder Diagnosis
by: Li, Yueyang, et al.
Published: (2025)
by: Li, Yueyang, et al.
Published: (2025)
LagKV: Lag-Relative Information of the KV Cache Tells Which Tokens Are Important
by: Liang, Manlai, et al.
Published: (2025)
by: Liang, Manlai, et al.
Published: (2025)
Learning Video Context as Interleaved Multimodal Sequences
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis
by: Wen, Yuhua, et al.
Published: (2025)
by: Wen, Yuhua, et al.
Published: (2025)
Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning
by: Luo, Meng, et al.
Published: (2026)
by: Luo, Meng, et al.
Published: (2026)
Wild2Avatar: Rendering Humans Behind Occlusions
by: Xiang, Tiange, et al.
Published: (2023)
by: Xiang, Tiange, et al.
Published: (2023)
Skeleton-in-Context: Unified Skeleton Sequence Modeling with In-Context Learning
by: Wang, Xinshun, et al.
Published: (2023)
by: Wang, Xinshun, et al.
Published: (2023)
Unveiling Context-Related Anomalies: Knowledge Graph Empowered Decoupling of Scene and Action for Human-Related Video Anomaly Detection
by: Chen, Chenglizhao, et al.
Published: (2024)
by: Chen, Chenglizhao, et al.
Published: (2024)
Leaving Some Facial Features Behind
by: Qiu, Cheng
Published: (2024)
by: Qiu, Cheng
Published: (2024)
On the Concept Trustworthiness in Concept Bottleneck Models
by: Huang, Qihan, et al.
Published: (2024)
by: Huang, Qihan, et al.
Published: (2024)
Similar Items
-
Representation Forcing for Bottleneck-Free Unified Multimodal Models
by: Wang, Yuqing, et al.
Published: (2026) -
Information Bottleneck Approach to Spatial Attention Learning
by: Lai, Qiuxia, et al.
Published: (2021) -
Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers
by: Hong, Jung-Ho, et al.
Published: (2025) -
Graph Integrated Multimodal Concept Bottleneck Model
by: Lin, Jiakai, et al.
Published: (2025) -
Unveiling the Visual Counting Bottleneck in Vision-Language Models
by: Pang, Xingzhou, et al.
Published: (2026)