Narrowing Information Bottleneck Theory for Multimodal Image-Text Representations Interpretability
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Zhiyu, Jin, Zhibo, Zhang, Jiayu, Yang, Nan, Huang, Jiahao, Zhou, Jianlong, Chen, Fang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attribution for Enhanced Explanation with Transferable Adversarial eXploration
by: Zhu, Zhiyu, et al.
Published: (2024)
by: Zhu, Zhiyu, et al.
Published: (2024)
Benchmarking Transferable Adversarial Attacks
by: Jin, Zhibo, et al.
Published: (2024)
by: Jin, Zhibo, et al.
Published: (2024)
Representation Forcing for Bottleneck-Free Unified Multimodal Models
by: Wang, Yuqing, et al.
Published: (2026)
by: Wang, Yuqing, et al.
Published: (2026)
Multimodal Conditional Information Bottleneck for Generalizable AI-Generated Image Detection
by: Qin, Haotian, et al.
Published: (2025)
by: Qin, Haotian, et al.
Published: (2025)
The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models
by: Serra, Alessandro Pietro, et al.
Published: (2024)
by: Serra, Alessandro Pietro, et al.
Published: (2024)
Spatial Information Bottleneck for Interpretable Visual Recognition
by: Shu, Kaixiang, et al.
Published: (2025)
by: Shu, Kaixiang, et al.
Published: (2025)
IBCapsNet: Information Bottleneck Capsule Network for Noise-Robust Representation Learning
by: Xiang, Canqun, et al.
Published: (2026)
by: Xiang, Canqun, et al.
Published: (2026)
An Empirical Study and Analysis of Text-to-Image Generation Using Large Language Model-Powered Textual Representation
by: Tan, Zhiyu, et al.
Published: (2024)
by: Tan, Zhiyu, et al.
Published: (2024)
GE-AdvGAN: Improving the transferability of adversarial samples by gradient editing-based adversarial generative model
by: Zhu, Zhiyu, et al.
Published: (2024)
by: Zhu, Zhiyu, et al.
Published: (2024)
Prototypical Information Bottlenecking and Disentangling for Multimodal Cancer Survival Prediction
by: Zhang, Yilan, et al.
Published: (2024)
by: Zhang, Yilan, et al.
Published: (2024)
Visual Explanations of Image-Text Representations via Multi-Modal Information Bottleneck Attribution
by: Wang, Ying, et al.
Published: (2023)
by: Wang, Ying, et al.
Published: (2023)
ReaSon: Reinforced Causal Search with Information Bottleneck for Video Understanding
by: Zhou, Yuan, et al.
Published: (2025)
by: Zhou, Yuan, et al.
Published: (2025)
FTII-Bench: A Comprehensive Multimodal Benchmark for Flow Text with Image Insertion
by: Ruan, Jiacheng, et al.
Published: (2024)
by: Ruan, Jiacheng, et al.
Published: (2024)
RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization
by: Huang, Mengqi, et al.
Published: (2024)
by: Huang, Mengqi, et al.
Published: (2024)
Concept Complement Bottleneck Model for Interpretable Medical Image Diagnosis
by: Wang, Hongmei, et al.
Published: (2024)
by: Wang, Hongmei, et al.
Published: (2024)
From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning
by: Luo, Ruilin, et al.
Published: (2026)
by: Luo, Ruilin, et al.
Published: (2026)
InfoBFR: Real-World Blind Face Restoration via Information Bottleneck
by: Gao, Nan, et al.
Published: (2025)
by: Gao, Nan, et al.
Published: (2025)
Learning Label-Efficient Interpretable Medical Image Diagnosis via Semi-supervised Hypergraph Concept Bottleneck Model
by: Yang, Yijun, et al.
Published: (2026)
by: Yang, Yijun, et al.
Published: (2026)
Towards SAR Automatic Target Recognition MultiCategory SAR Image Classification Based on Light Weight Vision Transformer
by: Zhao, Guibin, et al.
Published: (2024)
by: Zhao, Guibin, et al.
Published: (2024)
NOFT: Test-Time Noise Finetune via Information Bottleneck for Highly Correlated Asset Creation
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
by: Ma, Lichen, et al.
Published: (2026)
by: Ma, Lichen, et al.
Published: (2026)
Guarding the Gate: ConceptGuard Battles Concept-Level Backdoors in Concept Bottleneck Models
by: Lai, Songning, et al.
Published: (2024)
by: Lai, Songning, et al.
Published: (2024)
Information Bottleneck-Guided Heterogeneous Graph Learning for Interpretable Neurodevelopmental Disorder Diagnosis
by: Li, Yueyang, et al.
Published: (2025)
by: Li, Yueyang, et al.
Published: (2025)
Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers
by: Hong, Jung-Ho, et al.
Published: (2025)
by: Hong, Jung-Ho, et al.
Published: (2025)
Safety of Multimodal Large Language Models on Images and Texts
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
Multimodal Fusion via Self-Consistent Task-Gradient Fields
by: Xiong, Jiayu, et al.
Published: (2024)
by: Xiong, Jiayu, et al.
Published: (2024)
TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation
by: Wang, Xingrui, et al.
Published: (2024)
by: Wang, Xingrui, et al.
Published: (2024)
Pick-and-Draw: Training-free Semantic Guidance for Text-to-Image Personalization
by: Lv, Henglei, et al.
Published: (2024)
by: Lv, Henglei, et al.
Published: (2024)
TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
by: Luan, Bozhi, et al.
Published: (2024)
by: Luan, Bozhi, et al.
Published: (2024)
Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation
by: Zhao, Chenxi, et al.
Published: (2026)
by: Zhao, Chenxi, et al.
Published: (2026)
Open-set Cross Modal Generalization via Multimodal Unified Representation
by: Huang, Hai, et al.
Published: (2025)
by: Huang, Hai, et al.
Published: (2025)
The Semantic Lifecycle in Embodied AI: Acquisition, Representation and Storage via Foundation Models
by: Chen, Shuai, et al.
Published: (2026)
by: Chen, Shuai, et al.
Published: (2026)
EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models
by: Tan, Zhiyu, et al.
Published: (2024)
by: Tan, Zhiyu, et al.
Published: (2024)
Disentangled Representation Learning with Transmitted Information Bottleneck
by: Dang, Zhuohang, et al.
Published: (2023)
by: Dang, Zhuohang, et al.
Published: (2023)
PathAR: Structure-First Autoregressive Synthesis of Multimodal Pathology Images
by: Zhang, Yuan, et al.
Published: (2026)
by: Zhang, Yuan, et al.
Published: (2026)
Interpreting CLIP's Image Representation via Text-Based Decomposition
by: Gandelsman, Yossi, et al.
Published: (2023)
by: Gandelsman, Yossi, et al.
Published: (2023)
LEGO: Self-Supervised Representation Learning for Scene Text Images
by: Ren, Yujin, et al.
Published: (2024)
by: Ren, Yujin, et al.
Published: (2024)
Graph Information Bottleneck for Remote Sensing Segmentation
by: Shou, Yuntao, et al.
Published: (2023)
by: Shou, Yuntao, et al.
Published: (2023)
ELiTe: Efficient Image-to-LiDAR Knowledge Transfer for Semantic Segmentation
by: Zhang, Zhibo, et al.
Published: (2024)
by: Zhang, Zhibo, et al.
Published: (2024)
Learning Unsupervised Gaze Representation via Eye Mask Driven Information Bottleneck
by: Jiang, Yangzhou, et al.
Published: (2024)
by: Jiang, Yangzhou, et al.
Published: (2024)
Similar Items
-
Attribution for Enhanced Explanation with Transferable Adversarial eXploration
by: Zhu, Zhiyu, et al.
Published: (2024) -
Benchmarking Transferable Adversarial Attacks
by: Jin, Zhibo, et al.
Published: (2024) -
Representation Forcing for Bottleneck-Free Unified Multimodal Models
by: Wang, Yuqing, et al.
Published: (2026) -
Multimodal Conditional Information Bottleneck for Generalizable AI-Generated Image Detection
by: Qin, Haotian, et al.
Published: (2025) -
The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models
by: Serra, Alessandro Pietro, et al.
Published: (2024)